Insight · AI Solutions

    How to Evaluate AI Agents in Microsoft Foundry and Copilot Studio

    Agents fail quietly unless you test them. How to evaluate AI agents with Copilot Studio test sets and Foundry's rubric evaluators, traces-to-dataset and agent optimizer, and what to build first.

    Nick de Vrye, CTOPublished 8 October 20267 min read
    Navy Solv Systems title card reading 'How to Evaluate AI Agents' with a checklist motif.

    In Short: How to Evaluate AI Agents Before Users Do It for You

    Most agent projects test the happy path in a chat window and call it done. Then a model update, a new knowledge source or a reworded instruction changes the answers, and nobody notices until a user complains. Evaluating AI agents properly means a fixed set of test cases, clear scoring and a pass mark, re-run every time something changes.

    Both of Microsoft's agent platforms now have the tools for this. Copilot Studio has test sets with eight test methods. Microsoft Foundry has built-in agent evaluators, and in September it added rubric evaluators, datasets built from production traces and an agent optimizer. The tools differ because the builders differ, but the discipline is the same.

    Why Evaluation Is the Real Work

    An agent is a model, instructions, knowledge and tools working together. Change any one of them and the behaviour shifts. Models are replaced every few months: September alone brought GPT-6 and three new Claude models, as we cover in Claude in Microsoft Copilot and Foundry. Without a test set, every one of those changes is a guess.

    This is also the most common reason we see AI pilots stall. The demo worked. Nobody could prove the production version still did. An evaluation set turns "it seems fine" into a number that a business owner can sign off.

    Evaluating Agents in Copilot Studio

    Copilot Studio's agent evaluations became generally available in March 2026. Microsoft's documentation describes them for agents on the standard harness.

    Build a test set. Each test set holds up to 100 test cases. You can write them by hand, or import a CSV file with up to 100 questions of up to 1,000 characters each. You can also have AI generate questions from the agent's knowledge sources and topics, or build cases from themes in real user questions. Multi-turn conversation tests cover dialogue flows, not just single answers.

    Choose test methods. You can apply several to one test set:

    • General quality: an LLM judge checks relevance, groundedness, completeness and abstention. No expected answer needed
    • Compare meaning: intent similarity against an expected answer, with a default pass score of 50
    • Tool use: did the agent call the expected tools or topics
    • Keyword match, text similarity and exact match: for answers where the wording matters
    • Content safety: added in September 2026. It checks hate, sexual, violent and self-harm content against a severity threshold from 0 (strictest) to 7
    • Custom: your own evaluation instructions and pass/fail labels, such as "compliant" and "non-compliant" for an HR agent

    One trap: general quality can mark a correct refusal as a failure, because it checks whether the agent attempted an answer. If refusing is the right behaviour, test it with a custom method.

    Watch whose identity runs the test. Evaluations run as the signed-in user, against that user's knowledge and connections. Microsoft warns that generated test cases can include sensitive data that account can see. Use a test account whose access matches the real audience.

    Automate it. You can trigger evaluations through the Copilot Studio connector in Power Automate. A REST API through the Power Platform API is in preview for CI/CD pipelines. Version comparison then shows whether a change helped or caused a regression.

    Evaluating Agents in Microsoft Foundry

    Foundry gives engineers more control, and expects more effort in return.

    Built-in agent evaluators. Microsoft splits them into two groups. System evaluators judge the outcome: Task Completion, Task Adherence and Intent Resolution, all in preview. Process evaluators judge the tool calls: Tool Call Accuracy, Tool Input Accuracy, Tool Output Utilization and Tool Call Success. Two preview composites, Output Quality and Tool Use Quality, combine several checks into one judge call to cut cost and latency.

    Rubric evaluators. This is the part we would start with. A rubric is a set of weighted dimensions. An LLM judge scores each one from 1 to 5, and the weighted average becomes a 0-1 score with a default pass threshold of 0.5. Foundry can generate a first rubric from the agent's instructions, reference files and production traces. You then edit it until it separates good answers from bad ones. Microsoft recommends using rubrics "as your primary measure of agent quality", alongside built-in safety and groundedness evaluators.

    Traces to dataset. Foundry turns production traces into a versioned evaluation dataset. Set a cap from 1 to 1,000 samples and it picks a representative, de-duplicated subset. It redacts private content by default. Your test set then reflects what users actually ask.

    Agent optimizer. The optimizer runs your agent against a dataset, generates candidate instructions, tool descriptions and model choices (plus skills for hosted agents), and scores each candidate against the baseline. You decide whether to promote the winner. Microsoft's warning matters: any APIs or databases your agent calls really run during optimisation, so point it at test endpoints.

    On 24 September Microsoft said the rubric evaluator, traces-to-dataset generation and the agent optimizer would be generally available "later this month". Insights in Foundry, which finds recurring issues in production traces, is in public preview. Check the status label on each Learn page before you build a release process around it. Our Foundry Agent Service guide covers the platform these sit on.

    Which Platform's Evaluation Fits Your Agent

    Pick the platform for the builder and the agent, as in our Copilot Studio vs Foundry comparison. Each platform's evaluation tools then come with it:

    • Maker-built agents over documents and processes: Copilot Studio test sets, with general quality, compare meaning and a custom compliance test
    • Engineered agents with many tools: Foundry process evaluators plus a rubric, run from the SDK in your pipeline
    • High-volume production agents: traces to dataset, so the test set keeps up with real traffic, then the optimizer for model and instruction changes

    Evaluation is not free. Every run calls the agent and usually a judge model. Microsoft says agents on Copilot Studio's GitHub Copilot harness consume Copilot Credits for testing and evaluating as well as running. Size the test set per scenario, not per agent, and budget for it like any other test environment.

    The Release Gate We Use

    • Fifty to a hundred real cases per task type, from users or subject experts, not invented by the build team
    • A written pass mark agreed with the business owner before the first run
    • Re-run on every change: model, instructions, knowledge source or tool
    • Grounding checked separately from tone: many agents are polite and wrong, and bad source data shows up here first
    • Production traces fed back monthly, so the test set doesn't fall behind what users ask

    The same set does double duty for model choice. Our framework for choosing models in Foundry depends on it.

    Where Solv Systems Comes In

    Most of our time on agent projects goes into evaluation, not prompt writing. We build test sets from real cases, write the rubrics and custom tests with your subject experts, and wire the runs into your release process. If you have an agent in Copilot Studio or Foundry that nobody can vouch for, our AI automation consulting team can give it a baseline score as a short, fixed-scope piece of work.

    Sources and Further Reading

    Frequently asked

    Create a test set of up to 100 test cases, written by hand, imported from a CSV file, generated by AI from the agent's knowledge and topics, or built from themes in real user questions. Add one or more test methods, such as general quality, compare meaning, tool use or a custom test, then run it. Results can be compared across agent versions.

    A rubric evaluator scores agent responses against weighted criteria you define, using an LLM as the judge. Each dimension is scored from 1 to 5, and the weighted average is normalised to a 0-1 score with a default pass threshold of 0.5. Foundry can generate a first rubric from the agent's instructions, reference files and production traces.

    It runs your agent against an evaluation dataset, generates candidate changes to instructions, tool descriptions, model choice and, for hosted agents, skills, then scores each candidate against the baseline. You review the results and choose whether to promote the winner. Microsoft warns that any tools your agent calls really run during optimisation, so use test endpoints.

    Yes. Copilot Studio evaluations can be triggered through the Copilot Studio connector in Power Automate, and a REST API through the Power Platform API is in preview for CI/CD pipelines. Foundry evaluations run from the Python and JavaScript SDKs, which fit naturally into engineering pipelines.

    They can. Evaluations call the agent and usually an LLM judge, so they consume tokens or credits. Microsoft states that for Copilot Studio agents on the GitHub Copilot harness, usage-based billing applies to building, testing and evaluating agents. In Foundry, judge and optimisation models bill like any other model deployment.

    Microsoft does not set a minimum. In our experience, fifty to a hundred real cases per task type is enough to catch most regressions without making every run expensive. Copilot Studio caps a test set at 100 cases, which is a sensible size per scenario.