Agent Optimizer is currently in preview. This preview is provided without a service-level agreement, and we don’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
Supported agent types
The optimizer uses the same evaluation-driven improvement loop for both agent types, but the setup, optimization targets, and deployment experience differ.Prompt agents
For a prompt agent, the optimizer evaluates a selected agent version and generates alternative instructions. It can also improve descriptions for function-calling tools and compare the prompt across multiple candidate model deployments. Prompt-agent function-calling tools run on the client side. The optimizer can improve the tool and parameter descriptions that guide the model’s function calls, but it can’t evaluate the tool execution during tool-description optimization. Start the optimization wizard from the agent’s Optimize tab in the Foundry portal. The wizard guides you through selecting the agent version, dataset, evaluation criteria, and optimization models before you submit the run. You can generate a dataset from agent traces, select an existing Foundry dataset, or upload a dataset in the wizard. When the run finishes, compare score changes, review the before-and-after prompt, and inspect per-evaluator results. You can then promote a selected candidate to a new prompt-agent version from the optimization run. For the end-to-end portal experience, see Quickstart: Optimize a prompt agent.Hosted agents
For a hosted agent, the optimizer improves the configuration that your code loads at runtime. The agent needs the optimization package and a baseline configuration that exposes the instructions and, optionally, skills and function-calling tool definitions. Foundry Toolkit can automatically scaffold this integration. For the Azure Developer CLI workflow, add the integration before you run the optimizer. The optimizer can then generate and evaluate candidate configurations without changing your agent’s core logic. Hosted-agent optimization supports local JSONL datasets and datasets registered in your Foundry project, including evaluation datasets generated from traces. After a run, apply the selected candidate to your local configuration, review the changes, and redeploy the agent. For the end-to-end hosted-agent experience, see Quickstart: Optimize a hosted agent.The optimization workflow
Both agent types follow the same high-level path:- Select the agent baseline. Choose the prompt-agent version or hosted-agent configuration that you compare candidates against.
- Select an evaluation dataset. Use representative tasks from an existing or uploaded dataset. You can also use a registered dataset generated from agent traces.
- Select evaluation criteria. Choose built-in or custom evaluators that measure the behaviors you want to improve.
- Run the optimizer. Choose the available optimization targets and models, and then generate and evaluate candidates. For prompt agents, review the cost estimate in the portal before you submit the run.
- Review the results. Compare candidate scores against your baseline, and then pick the best candidate. When available, review the post-run measured token usage. See Understand optimization results.
- Apply the selected candidate. Promote a prompt-agent candidate to a new agent version, or apply the hosted-agent configuration and redeploy.
How the agent optimizer works
The agent optimizer runs a closed-loop evaluation and improvement cycle:- Evaluate the baseline. The optimizer invokes your agent against a dataset of tasks and scores each response against criteria you define or a built-in default set. The baseline is your agent’s score before any changes.
- Generate candidates. The optimizer produces alternative configurations called candidates, such as rewritten instructions, discovered skills, improved tool descriptions, or different model selections. Available candidate types depend on the agent type.
- Evaluate candidates. The optimizer tests each candidate against the same dataset.
- Rank and recommend. The optimizer ranks results by composite score, a value between 0.0 and 1.0 that represents aggregate performance, and marks the best candidate with ★. You then apply and deploy the winner.
load_config() returns your baseline normally and supplies optimized configuration during a run without feature flags or conditional logic.
Optimization targets
An optimization target is a specific aspect of your agent’s configuration that the optimizer can improve. Available targets depend on the agent type. For hosted agents, the optimizer automatically activates applicable targets from the baseline configuration andeval.yaml settings.
For hosted-agent baseline inputs, see Make your agent optimizer-ready. To run and configure hosted-agent targets, see Optimize agent instructions, skills, tools, and models.
Models
The agent optimizer uses two models during an optimization run. Both must be deployed in your Foundry project.
The eval model runs once for each evaluator, task, and candidate evaluation. It
reads the agent’s response and each criterion, then returns a binary score. The
optimization model analyzes baseline results and generates improved candidates
across the configured targets, including instructions, skills, tools, and
models. Because it reasons over the full dataset, a more capable optimization
model typically produces better candidates.
For prompt agents, select the eval model and candidate models in the optimization wizard. For hosted agents, specify the models in
eval.yaml or with CLI flags. The optimization_model setting is required for hosted-agent runs. For configuration steps, see Choose the eval and optimization models.
Understand optimization results
This section explains the results table, how the composite score is computed, and how to interpret improvements. For prompt agents, a completed run shows the baseline and candidate scores, a before-and-after prompt, model, and tools description comparison, and per-evaluator scores. For hosted agents, the CLI also displays a results table:Results table columns
The ★ marks the candidate with the highest composite score. This is the recommended candidate to deploy.
How scores are computed
Each evaluator in your dataset produces a raw score for the agent’s response. The optimizer processes these scores to produce the composite score shown in results:- Rescale: Each evaluator’s raw score is rescaled to 0–1.
- Flip if needed: If an evaluator is configured so that lower is better, the score is flipped so that all evaluators use “higher is better” semantics.
- Average: The rescaled scores across all evaluators and tasks are averaged to produce the composite score.
Interpret score improvements
Token trade-offs
Optimized instructions are often longer and more detailed, which can increase response token usage. Consider these factors:- Whether the token increase is proportional to the score improvement
- Whether the cost increase fits your budget
- Whether responses are unnecessarily verbose or adding value with the extra length
Limitations and availability
- The agent optimizer supports prompt agents and hosted agents during preview.
- Prompt-agent optimization runs start in the Foundry portal. You can promote a selected candidate to a new agent version from the completed run.
- Hosted-agent optimization is available in all regions where hosted agents are available, except Norway East.
- Hosted-agent optimization requires the Responses protocol.
Related content
- Quickstart: Create a prompt agent
- Quickstart: Optimize a prompt agent
- Quickstart: Optimize a hosted agent
- Agent optimizer cost estimates and token usage
- Make your agent optimizer-ready
- Create an evaluation dataset and evaluators
- Optimize agent instructions, skills, tools, and models
- Convert agent traces into evaluation datasets
- Run agent evaluations with the azd CLI