Items marked preview in this article are currently in preview. This preview is provided without a service-level agreement, and Microsoft doesn’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
- System evaluation - to examine the end-to-end outcomes of the agentic system.
- Process evaluation - to verify the step-by-step execution to achieve the outcomes.
System evaluation
System evaluation examines the quality of the final outcome of your agentic workflow. These evaluators are applicable to single agents and, in multi-agent systems, to the main orchestrator or the final agent responsible for task completion:- Task Completion - Did the agent fully complete the requested task?
- Customer Satisfaction - How satisfied would a user be with the agent’s performance?
- Task Adherence - Did the agent follow the rules and constraints in its instructions?
- Task Navigation Efficiency - Did the agent perform the expected steps efficiently?
- Intent Resolution - Did the agent correctly identify and address user intentions?
Relevance and Groundedness that take agentic inputs to assess the final response quality.
Examples:
- Task completion (preview) sample
- Task adherence sample
- Task navigation efficiency sample
- Intent resolution sample
Process evaluation
Process evaluation examines the quality and efficiency of each step in your agentic workflow. These evaluators focus on the tool calls executed in a system to complete tasks:- Tool Call Accuracy - Did the agent make the right tool calls with correct parameters without redundancy?
- Tool Selection - Did the agent select the correct and necessary tools?
- Tool Input Accuracy - Did the agent provide correct parameters for tool calls?
- Tool Output Utilization - Did the agent correctly use tool call results in its reasoning and final response?
- Tool Call Success - Did the tool calls succeed without technical errors?
- Tool call accuracy sample
- Tool selection sample
- Tool input accuracy sample
- Tool output utilization sample
- Tool call success sample
Quality evaluation (preview)
Quality evaluation assesses the overall quality of an AI assistant’s response at the turn level. The Quality Grader evaluator is the same quality evaluator used in Microsoft Copilot Studio agent evaluation. It examines multiple dimensions of response quality:- Relevance - Is the response relevant to the user’s query?
- Abstention - Does the agent appropriately abstain when it cannot or should not answer?
- Answer completeness - Does the response fully address the user’s question?
- Groundedness - Is the response grounded in the provided context?
- Context coverage - Does the response make use of the relevant information in the context?
Composite evaluators (preview)
Composite evaluators measure several quality dimensions in one LLM judge call. Use them to reduce the cost and latency of running the corresponding built-in evaluators separately while retaining a score and reason for each quality dimension.Use a model from the GPT-5.6 family as the LLM judge for composite evaluators.
For the lowest-cost option that maintains high evaluation quality, use
gpt-5.6-luna.Output Quality
The Output Quality evaluator (builtin.output_quality) assesses the quality of
an agent’s response and its success in addressing the user’s task. It batches
the following evaluators into one LLM judge call:
Tool Use Quality
The Tool Use Quality evaluator (builtin.tool_use_quality) assesses an agent’s
tool-use process from tool selection through use of the returned result. It
batches the following evaluators into one LLM judge call:
Composite results
The primaryoutput_quality or tool_use_quality result passes only when every
applicable component passes. A component that isn’t applicable is skipped and
doesn’t cause the primary result to fail. If all components are skipped, the
primary result is not_applicable.
Each composite evaluator returns a primary binary result and the results of its
component evaluators. Use the primary result for a strict quality gate, and use
the component results to identify the quality dimension that caused a failure.
For example, an Output Quality result fails when Fluency, Coherence, Intent
Resolution, Task Adherence, and Task Completion pass but Groundedness fails.
The
output_quality_reason identifies Groundedness as the failed component, and
the groundedness_reason explains its score.
Configure composite evaluators
Both composite evaluators support turn-level and conversation-level evaluation. The input shape determines the evaluation level by default:- Map
queryandresponsefor turn-level evaluation. - Map
messagesfor conversation-level evaluation.
evaluation_level initialization parameter to turn or conversation
to override this behavior.
The following configuration runs both composite evaluators at the turn level.
The response contains agent messages with tool calls and tool results, and the
tool definitions describe the tools available to the agent.
messages
and either omit evaluation_level or set it to conversation:
Model and tool support
For AI-assisted evaluators, you can use Azure OpenAI or OpenAI reasoning models and non-reasoning models for the LLM judge.Supported tools
Agent evaluators support the following tools:- File Search
- Function Tool (user-defined tools)
- MCP
- Knowledge-based MCP
tool_call_accuracy, tool input accuracy, tool_output_utilization, tool_call_success, or groundedness evaluators if your agent conversation includes calls to these tools:
- Azure AI Search
- Bing Grounding
- Bing Custom Search
- SharePoint Grounding
- Code Interpreter
- Fabric Data Agent
- Web Search
Using agent evaluators
Agent evaluators assess how well AI agents perform tasks, follow instructions, and use tools effectively. Each evaluator requires specific data mappings and parameters:
Turn-level evaluation scores an individual agent response; conversation-level evaluation scores the full interaction. Use
evaluation_level to select the level when the evaluator supports both. For the message structure, see Messages with tool calls.
Example input
Your test dataset should contain the fields referenced in your data mappings. Examples for the format below:query and response. These arrays use the same OpenAI message structure as the preferred messages schema. See Separate query and response format. The system message is optional but useful for evaluators that assess agent behavior against instructions, including task_adherence, task_completion, tool_call_accuracy, tool_selection, tool_input_accuracy, tool_output_utilization, and groundedness:
Tool definitions format
Thetool_definitions field describes the tools available to the agent. It contains a list of tool objects with a name, description, and JSON Schema parameters object:
tool_definitions field in your test dataset alongside query and response.
Configuration example
Data mapping syntax:{{item.field_name}}references fields from your test dataset (for example,{{item.query}}).{{sample.output_items}}references the agent’s structured output, including tool calls and results. Use this for evaluators that need full interaction context (task_adherence,tool_call_accuracy,tool_selection,tool_input_accuracy,tool_output_utilization).{{sample.output_text}}references the agent’s plain text response. Use this for evaluators that expect a string response (for example,coherence,violence).
Example output
Agent evaluators return Pass/Fail results with reasoning. Key output fields:intent_resolution and tool_call_accuracy), the output includes a numeric score field alongside the pass/fail result:
Task navigation efficiency
Task Navigation Efficiency measures whether the agent took an optimal sequence of actions by comparing against an expected sequence (ground truth). Use this evaluator for workflow optimization and regression testing.
Actions format:
Map either
actions or messages with expected_actions. The actions field accepts a string or a list of message objects. The messages field accepts the full interaction as an array of message objects. Each action message represents a step the agent took during the conversation:
The
actions and expected_actions fields use different formats. actions contains the agent’s actual behavior as text or message objects, while expected_actions contains the expected tool names and, optionally, their parameters.actions, map messages and expected_actions:
expected_actions can be a simple list of expected steps: