This feature is currently in public preview. This preview is provided without a service-level agreement, and we don’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
How tools work
When you type a natural-language prompt in an MCP-compliant client (for example, GitHub Copilot Agent Mode), the language model selects the appropriate tool and formulates the required parameters on your behalf. You don’t call tools directly — you describe what you want, and the model translates your intent into a tool call. Each tool is classified as read (retrieves information) or write (creates, updates, or deletes resources). Write operations affect live resources and billing immediately. Review the security best practices before running write operations.Permissions
All operations run with the authenticated user’s permissions through the Microsoft Entra ID On-Behalf-Of flow. You need the following roles:
For more information, see Role-based access control for Microsoft Foundry.
Key identifiers
Many tools require resource identifiers. The language model extracts these from your prompt context, but it helps to know the formats:Agent management
Manage the full lifecycle of agents in a Foundry project, including creation, invocation, container orchestration, and deletion. Example prompts:- “List all agents in my Foundry project.”
- “Create a new agent named
faq-agentusing modelgpt-4o-mini.” - “Send ‘Hello, how can you help?’ to my
customer-support-agent.” - “Start the container for my hosted agent
triage-agent.” - “Check the container status for
triage-agent.” - “Show me the agent definition schema for prompt agents.”
- “Delete the
old-test-agentfrom my project.”
Hosted agent sessions and files
Create and manage isolated compute sessions for hosted agents. Inspect logs and manage files in each session’s filesystem. Example prompts:- “Create a session for my hosted agent
data-analyst.” - “List the sessions for
data-analystand show the newest first.” - “Stream up to 100 log lines from session
analysis-2026.” - “Upload
input.csvto/data/input.csvin the session.” - “Download
/data/results.csvfrom the session.” - “Delete session
analysis-2026and release its compute resources.”
Toolbox management
Create, retrieve, version, update, and delete toolboxes in a Foundry project. Example prompts:- “Show me the
customer-support-toolstoolbox.” - “Get version 2 of
customer-support-tools.” - “Create a new version of
customer-support-tools.” - “Set version 2 of
customer-support-toolsas the default.” - “Set version 1 of
customer-support-toolsas the default, then delete version 2.” - “Delete the
old-support-toolstoolbox.”
toolbox_version_create also creates the toolbox if it doesn’t exist. For an existing toolbox, creating a version doesn’t change the default version. Call toolbox_update with defaultVersion set to the new version to promote it. Before you delete the current default version, set another version as the default.
Dataset management
Create, retrieve, version, and download evaluation datasets in a Foundry project. Example prompts:- “Upload my customer support Q&A dataset from this Azure Blob Storage URL.”
- “Show me all datasets in my Foundry project.”
- “Get details for the
customer-support-qadataset version 2.” - “List all versions of my
product-reviewsdataset.” - “Get a download URL for the latest version of
customer-support-qa.”
Data generation jobs
Create and manage standalone data generation jobs for evaluation or fine-tuning data. Generate data from prompts, agents, traces, datasets, or uploaded files. Example prompts:- “Generate 100 evaluation question-and-answer samples from this prompt.”
- “Show me the status of data generation job
job-123.” - “Cancel data generation job
job-123.” - “Delete the completed data generation job
job-123.”
Evaluation operations
Run batch evaluations against agents or datasets, and compare results across runs. Example prompts:- “Evaluate my
customer-support-agentv2 using Relevance, Groundedness, and Coherence evaluators.” - “Run a batch evaluation on my JSONL dataset with Violence and HateUnfairness evaluators.”
- “Generate 50 synthetic test queries and evaluate my agent with them.”
- “Show me all evaluation runs in my Foundry project.”
- “Compare run-baseline-123 against treatment runs run-124 and run-125.”
Trace evaluation
Run multi-turn conversation evaluations over agent traces or explicitly selected conversations and W3C trace IDs. Example prompts:- “Evaluate up to five traces from
customer-support-agentfrom the last 24 hours for groundedness and task completion.” - “Evaluate these conversation IDs for customer satisfaction at the conversation level.”
- “Evaluate these W3C trace IDs with my custom
tone-checkevaluator.”
Evaluation suites
Create, version, run, update, and delete evaluation suites that group datasets, target agents, and evaluators. You can also generate suites from prompts, agents, traces, datasets, or uploaded files. Example prompts:- “Create an evaluation suite named
support-qualityfor my support dataset and agent.” - “Run version 2 of
support-quality.” - “Update the display name and tags for version 2 of
support-quality.” - “Generate an evaluation suite from recent traces for
customer-support-agent.” - “Show me all evaluation suite generation jobs.”
Evaluator catalog
Browse built-in evaluators and manage custom evaluators for use in evaluation runs. Example prompts:- “List all built-in evaluators available in my project.”
- “Show me the full definition of the
coherenceevaluator.” - “Create a custom prompt-based evaluator called
tone-checkthat scores responses on a 1-5 scale.” - “Update the description of my
tone-checkevaluator.” - “Delete version 1 of my
old-evaluator.”
Evaluator generation jobs
Generate evaluators adaptively from prompts, agents, or datasets, and monitor or cancel the generation jobs. Example prompts:- “Generate an evaluator named
policy-compliancefrom this policy prompt.” - “Regenerate
policy-compliancefrom version 2 of my support dataset using model deploymentgpt-4o.” - “Show me all running evaluator generation jobs.”
- “Cancel evaluator generation job
job-456.”
Model catalog and details
Explore models in the Foundry model catalog, get model details, and determine whether a model can be deployed to Managed Compute. Example prompts:- “Show me all GPT-5.4 models available in the catalog.”
- “List all Microsoft-published models with MIT license.”
- “Get detailed information and code samples for GPT-5-mini.”
- “Show me models advertised for Managed Compute.”
- “Get Managed Compute details and deployment templates for
openai--gpt-oss-20b.”
isManagedCompute in model_catalog_list identifies an advertised candidate. Confirm deployability with model_details_get: isFoundryManagedCompute is the authoritative signal, and the resolved deployment templates identify the available deployment options and accelerators.
Model deployment management
Deploy, inspect, and remove model deployments, including Managed Compute deployments, in a Foundry account. Example prompts:- “Use
model_deploy_azure_direct_modelto deploy GPT-5-mini asproduction-chatbotwith 20 capacity units.” - “Show me all my current model deployments.”
- “Show me all my Managed Compute deployments.”
- “Use
model_deploy_managed_computeto deployopenai--gpt-oss-20basgpt-oss-managedusing its A100_80GB deployment template.” - “Delete the
old-test-deploymentthat I’m no longer using.”
Don’t use the deprecated
model_deploy alias for new calls. For Managed Compute, first use model_catalog_list as a candidate signal, then call model_details_get. Use model_deploy_managed_compute only when isFoundryManagedCompute is true, and pass one of the resolved deployment templates. Use model_deployment_get to retrieve either deployment type, and model_deployment_delete to delete either type.
Model analytics and recommendations
Compare model benchmarks and get recommendations for switching to more cost-effective or higher-quality models. Example prompts:- “Show me benchmark data for all available models.”
- “Compare benchmark performance between GPT-5.4 and GPT-4.”
- “Find models similar to my current GPT-4 deployment.”
- “What models would give me better quality/cost ratio than what I’m using now?”
Model monitoring and operations
Track deployment health, monitor metrics, check deprecation status, and view quota usage, including Managed Compute accelerator quota. Example prompts:- “Show me the request metrics for my
production-chatbotdeployment.” - “Check if any of my deployments are using deprecated model versions.”
- “Show me quota usage across all regions for my subscription.”
- “Show me A100_80GB, H100_80GB, MI300_192GB, and H200_141GB accelerator quota in East US 2.”
Project connections
Manage connections to external services (Azure OpenAI, Azure Blob Storage, search, and others) within a Foundry project. Example prompts:- “List all connections in my Foundry project.”
- “Show me the details for my
azure-searchconnection.” - “What connection types and authentication methods are supported?”
- “Create a new AzureOpenAI connection called
my-openaiusing AAD auth.” - “Delete the
old-storageconnection from my project.”
Prompt optimization
Optimize system prompts and developer messages for better LLM performance. Example prompts:- “Optimize my system prompt: ‘You are a helpful customer service agent’ using
gpt-5.4.” - “Improve my agent instructions to get more concise responses.”
- “Refine my optimized prompt to also handle follow-up questions.”
Continuous evaluation
Enable, monitor, and manage continuous evaluation for agents. Continuous evaluation automatically evaluates agent responses on an ongoing basis. The tool auto-detects the agent kind (prompt or hosted) and configures the appropriate evaluation mechanism — evaluation rules for prompt agents, or scheduled evaluation runs for hosted agents. Example prompts:- “Enable continuous evaluation for my
customer-support-agentusing Relevance and Groundedness evaluators.” - “Set up continuous evaluation for my hosted agent
triage-agentto run every 2 hours.” - “Show me the continuous evaluation configuration for
customer-support-agent.” - “Disable continuous evaluation for my
old-test-agent.” - “Update continuous evaluation for
customer-support-agentto sample 50% of responses.”
Example workflows
Model deployment and optimization:- “Show me the
gpt-5.6-solmodel in the catalog.” - “Deploy
gpt-5.6-solascustomer-service-botwith 15 capacity units.” - “Monitor the request latency for my new deployment.”
- “Recommend more cost-effective alternatives based on current usage.”
- “Use
model_quota_listto show my Managed Compute accelerator quota and usage for subscription<subscription-id>in East US 2.” - “Compare the available instances for
A100_80GB,H100_80GB,MI300_192GB, andH200_141GB.” - “I need four instances. Identify which accelerators currently have enough available quota, and remind me that quota availability doesn’t guarantee deployment capacity.”
- “Before any deployment, show me the selected accelerator and instance count for review because Managed Compute allocates billable dedicated GPU capacity.”
- “Use
model_catalog_listto find models advertised for Managed Compute, whereisManagedComputeistrue.” - “For
openai--gpt-oss-20b, usemodel_details_getto confirm thatisFoundryManagedComputeistrue, and show its registrymodelAssetId, resolved deployment templates, and accelerators.” - “Use
model_quota_listto check the selected accelerator in East US 2 before deployment.” - “Review the deployment template, accelerator, and instance count with me. After I confirm the billable dedicated GPU allocation, use
model_deploy_managed_computeto creategpt-oss-managed-testwith the registrymodelAssetId, one resolveddeploymentTemplate, and the approvedinstanceCount.” - “Use
model_deployment_getto monitorgpt-oss-managed-testand verify that itsKindisManaged.” - “When testing is complete and I confirm cleanup, use
model_deployment_deleteto deletegpt-oss-managed-test.”
- “List all agents in my project.”
- “Evaluate my
customer-support-agentv2 using Relevance, Groundedness, and Safety evaluators.” - “Compare my baseline evaluation against the new run.”
- “Show me the comparison results with statistical significance.”
- “List all agents in my project.”
- “Enable continuous evaluation for my
customer-support-agentusing Relevance, Groundedness, and Safety evaluators with the existinggpt-5.6-lunadeployment.” - “Show me the continuous evaluation configuration for
customer-support-agent.” - “Update continuous evaluation to sample 25% of responses with a maximum of 10 runs per hour.”
- “List all my current deployments and their usage.”
- “Check which deployments are using deprecated model versions.”
- “Show me my quota usage across all regions.”
- “Delete unused test deployments to free up capacity.”
Preview limitations
Foundry MCP Server is in public preview. The following limitations apply:- No network isolation — Foundry MCP Server uses the public endpoint
https://mcp.ai.azure.com. Resources behind Azure Private Links aren’t accessible. For private MCP connectivity, build your own MCP server and connect it to Agent Service with private networking. - Data residency — Requests and responses might be processed in EU or US data centers. The server itself doesn’t store data, but cross-region processing can occur.
- No SLA — Preview features don’t include a service-level agreement. Don’t use the server for production workloads that require guaranteed availability.
- Tool set might change — Tools, parameters, and return values might change during the preview period without notice.
Common errors
For more troubleshooting guidance, see Foundry MCP Server security and best practices.
Related content
- Get started with Foundry MCP Server
- Learn how to build your own MCP server
- Review security best practices for MCP servers
- Learn how to submit your MCP server for certification at Microsoft MCP server certification overview.