Run an evaluation for an agent (preview)

[This article is prerelease documentation and is subject to change.]

After you create a test set with conversations, run an evaluation to measure your agent's performance. The evaluation processes each conversation and produces scored results based on the selected test method.

Important

Prerequisites

To run an evaluation for an agent:

  1. Open your agent in Copilot Studio.
  2. Select the Evaluate tab.
  3. In the Evaluation dropdown, select the evaluation you want to run.
  4. Add at least one conversation to the test set if you haven't already.
  5. In the Configure test set panel, verify:
    • The evaluation Name is set.
    • The Test method is configured (for example, General quality).
  6. Select Evaluate to start the evaluation. The evaluation processes each conversation in the test set. Depending on the number of conversations, this process might take several minutes.

Tip

Run the same evaluation multiple times. Each run is saved separately, so you can compare results across runs and see how changes to your agent affect quality.

Re-run an evaluation

After you change your agent's instructions, knowledge, or tools, re-run the evaluation to measure the impact:

  1. On the Evaluate tab, select the evaluation you want to re-run from the Evaluation dropdown.
  2. Under Recent results, select Evaluate test set again or select the Evaluate test set icon within the test set to start a new run.
  3. Compare the new run's results with previous runs to see if the changes improved or degraded performance.

Download an evaluation session

When the Evaluate tab is available for your environment, you can download session data from an evaluation run to review it offline or share it with a reviewer. The download packages the evaluation run's metadata (evaluation name, run identifier, agent version, and test method) together with the per-conversation results (each turn's input, agent response, and score).

To download an evaluation session:

  1. In the Evaluate tab, open the evaluation whose run you want to download.
  2. Under Recent results, select the run you want to download.
  3. Select Download to save the session data.

Use downloaded session data to:

  • Review agent responses offline when you don't have access to Copilot Studio.
  • Share a specific run with a reviewer for feedback.
  • Keep a record of an evaluation run alongside your test artifacts.

Downloaded session data reflects the state of the run at the time you download it. Re-run the evaluation if you want to capture the effect of later changes to the agent.