Skip to main content
This reference documents 79 Foundry MCP Server tools across 17 categories that let you manage agents, sessions, toolboxes, datasets, evaluations, model deployments, continuous evaluation, and more — all through conversational prompts instead of API calls. Use it to explore each tool and try the example prompts in your own project.
Before using these tools, complete the Foundry MCP Server setup.
This feature is currently in public preview. This preview is provided without a service-level agreement, and we don’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.

How tools work

When you type a natural-language prompt in an MCP-compliant client (for example, GitHub Copilot Agent Mode), the language model selects the appropriate tool and formulates the required parameters on your behalf. You don’t call tools directly — you describe what you want, and the model translates your intent into a tool call. Each tool is classified as read (retrieves information) or write (creates, updates, or deletes resources). Write operations affect live resources and billing immediately. Review the security best practices before running write operations.

Permissions

All operations run with the authenticated user’s permissions through the Microsoft Entra ID On-Behalf-Of flow. You need the following roles: For more information, see Role-based access control for Microsoft Foundry.

Key identifiers

Many tools require resource identifiers. The language model extracts these from your prompt context, but it helps to know the formats:

Agent management

Manage the full lifecycle of agents in a Foundry project, including creation, invocation, container orchestration, and deletion. Example prompts:
  • “List all agents in my Foundry project.”
  • “Create a new agent named faq-agent using model gpt-4o-mini.”
  • “Send ‘Hello, how can you help?’ to my customer-support-agent.”
  • “Start the container for my hosted agent triage-agent.”
  • “Check the container status for triage-agent.”
  • “Show me the agent definition schema for prompt agents.”
  • “Delete the old-test-agent from my project.”

Hosted agent sessions and files

Create and manage isolated compute sessions for hosted agents. Inspect logs and manage files in each session’s filesystem. Example prompts:
  • “Create a session for my hosted agent data-analyst.”
  • “List the sessions for data-analyst and show the newest first.”
  • “Stream up to 100 log lines from session analysis-2026.”
  • “Upload input.csv to /data/input.csv in the session.”
  • “Download /data/results.csv from the session.”
  • “Delete session analysis-2026 and release its compute resources.”

Toolbox management

Create, retrieve, version, update, and delete toolboxes in a Foundry project. Example prompts:
  • “Show me the customer-support-tools toolbox.”
  • “Get version 2 of customer-support-tools.”
  • “Create a new version of customer-support-tools.”
  • “Set version 2 of customer-support-tools as the default.”
  • “Set version 1 of customer-support-tools as the default, then delete version 2.”
  • “Delete the old-support-tools toolbox.”
Toolbox versions are immutable. toolbox_version_create also creates the toolbox if it doesn’t exist. For an existing toolbox, creating a version doesn’t change the default version. Call toolbox_update with defaultVersion set to the new version to promote it. Before you delete the current default version, set another version as the default.

Dataset management

Create, retrieve, version, and download evaluation datasets in a Foundry project. Example prompts:
  • “Upload my customer support Q&A dataset from this Azure Blob Storage URL.”
  • “Show me all datasets in my Foundry project.”
  • “Get details for the customer-support-qa dataset version 2.”
  • “List all versions of my product-reviews dataset.”
  • “Get a download URL for the latest version of customer-support-qa.”

Data generation jobs

Create and manage standalone data generation jobs for evaluation or fine-tuning data. Generate data from prompts, agents, traces, datasets, or uploaded files. Example prompts:
  • “Generate 100 evaluation question-and-answer samples from this prompt.”
  • “Show me the status of data generation job job-123.”
  • “Cancel data generation job job-123.”
  • “Delete the completed data generation job job-123.”

Evaluation operations

Run batch evaluations against agents or datasets, and compare results across runs. Example prompts:
  • “Evaluate my customer-support-agent v2 using Relevance, Groundedness, and Coherence evaluators.”
  • “Run a batch evaluation on my JSONL dataset with Violence and HateUnfairness evaluators.”
  • “Generate 50 synthetic test queries and evaluate my agent with them.”
  • “Show me all evaluation runs in my Foundry project.”
  • “Compare run-baseline-123 against treatment runs run-124 and run-125.”

Trace evaluation

Run multi-turn conversation evaluations over agent traces or explicitly selected conversations and W3C trace IDs. Example prompts:
  • “Evaluate up to five traces from customer-support-agent from the last 24 hours for groundedness and task completion.”
  • “Evaluate these conversation IDs for customer satisfaction at the conversation level.”
  • “Evaluate these W3C trace IDs with my custom tone-check evaluator.”

Evaluation suites

Create, version, run, update, and delete evaluation suites that group datasets, target agents, and evaluators. You can also generate suites from prompts, agents, traces, datasets, or uploaded files. Example prompts:
  • “Create an evaluation suite named support-quality for my support dataset and agent.”
  • “Run version 2 of support-quality.”
  • “Update the display name and tags for version 2 of support-quality.”
  • “Generate an evaluation suite from recent traces for customer-support-agent.”
  • “Show me all evaluation suite generation jobs.”

Evaluator catalog

Browse built-in evaluators and manage custom evaluators for use in evaluation runs. Example prompts:
  • “List all built-in evaluators available in my project.”
  • “Show me the full definition of the coherence evaluator.”
  • “Create a custom prompt-based evaluator called tone-check that scores responses on a 1-5 scale.”
  • “Update the description of my tone-check evaluator.”
  • “Delete version 1 of my old-evaluator.”

Evaluator generation jobs

Generate evaluators adaptively from prompts, agents, or datasets, and monitor or cancel the generation jobs. Example prompts:
  • “Generate an evaluator named policy-compliance from this policy prompt.”
  • “Regenerate policy-compliance from version 2 of my support dataset using model deployment gpt-4o.”
  • “Show me all running evaluator generation jobs.”
  • “Cancel evaluator generation job job-456.”

Model catalog and details

Explore models in the Foundry model catalog, get model details, and determine whether a model can be deployed to Managed Compute. Example prompts:
  • “Show me all GPT-5.4 models available in the catalog.”
  • “List all Microsoft-published models with MIT license.”
  • “Get detailed information and code samples for GPT-5-mini.”
  • “Show me models advertised for Managed Compute.”
  • “Get Managed Compute details and deployment templates for openai--gpt-oss-20b.”
For Managed Compute, isManagedCompute in model_catalog_list identifies an advertised candidate. Confirm deployability with model_details_get: isFoundryManagedCompute is the authoritative signal, and the resolved deployment templates identify the available deployment options and accelerators.

Model deployment management

Deploy, inspect, and remove model deployments, including Managed Compute deployments, in a Foundry account. Example prompts:
  • “Use model_deploy_azure_direct_model to deploy GPT-5-mini as production-chatbot with 20 capacity units.”
  • “Show me all my current model deployments.”
  • “Show me all my Managed Compute deployments.”
  • “Use model_deploy_managed_compute to deploy openai--gpt-oss-20b as gpt-oss-managed using its A100_80GB deployment template.”
  • “Delete the old-test-deployment that I’m no longer using.”
Choose explicit tools for each deployment type: Don’t use the deprecated model_deploy alias for new calls. For Managed Compute, first use model_catalog_list as a candidate signal, then call model_details_get. Use model_deploy_managed_compute only when isFoundryManagedCompute is true, and pass one of the resolved deployment templates. Use model_deployment_get to retrieve either deployment type, and model_deployment_delete to delete either type.

Model analytics and recommendations

Compare model benchmarks and get recommendations for switching to more cost-effective or higher-quality models. Example prompts:
  • “Show me benchmark data for all available models.”
  • “Compare benchmark performance between GPT-5.4 and GPT-4.”
  • “Find models similar to my current GPT-4 deployment.”
  • “What models would give me better quality/cost ratio than what I’m using now?”

Model monitoring and operations

Track deployment health, monitor metrics, check deprecation status, and view quota usage, including Managed Compute accelerator quota. Example prompts:
  • “Show me the request metrics for my production-chatbot deployment.”
  • “Check if any of my deployments are using deprecated model versions.”
  • “Show me quota usage across all regions for my subscription.”
  • “Show me A100_80GB, H100_80GB, MI300_192GB, and H200_141GB accelerator quota in East US 2.”

Project connections

Manage connections to external services (Azure OpenAI, Azure Blob Storage, search, and others) within a Foundry project. Example prompts:
  • “List all connections in my Foundry project.”
  • “Show me the details for my azure-search connection.”
  • “What connection types and authentication methods are supported?”
  • “Create a new AzureOpenAI connection called my-openai using AAD auth.”
  • “Delete the old-storage connection from my project.”

Prompt optimization

Optimize system prompts and developer messages for better LLM performance. Example prompts:
  • “Optimize my system prompt: ‘You are a helpful customer service agent’ using gpt-5.4.”
  • “Improve my agent instructions to get more concise responses.”
  • “Refine my optimized prompt to also handle follow-up questions.”

Continuous evaluation

Enable, monitor, and manage continuous evaluation for agents. Continuous evaluation automatically evaluates agent responses on an ongoing basis. The tool auto-detects the agent kind (prompt or hosted) and configures the appropriate evaluation mechanism — evaluation rules for prompt agents, or scheduled evaluation runs for hosted agents. Example prompts:
  • “Enable continuous evaluation for my customer-support-agent using Relevance and Groundedness evaluators.”
  • “Set up continuous evaluation for my hosted agent triage-agent to run every 2 hours.”
  • “Show me the continuous evaluation configuration for customer-support-agent.”
  • “Disable continuous evaluation for my old-test-agent.”
  • “Update continuous evaluation for customer-support-agent to sample 50% of responses.”

Example workflows

Model deployment and optimization:
  • “Show me the gpt-5.6-sol model in the catalog.”
  • “Deploy gpt-5.6-sol as customer-service-bot with 15 capacity units.”
  • “Monitor the request latency for my new deployment.”
  • “Recommend more cost-effective alternatives based on current usage.”
Managed Compute quota planning:
  • “Use model_quota_list to show my Managed Compute accelerator quota and usage for subscription <subscription-id> in East US 2.”
  • “Compare the available instances for A100_80GB, H100_80GB, MI300_192GB, and H200_141GB.”
  • “I need four instances. Identify which accelerators currently have enough available quota, and remind me that quota availability doesn’t guarantee deployment capacity.”
  • “Before any deployment, show me the selected accelerator and instance count for review because Managed Compute allocates billable dedicated GPU capacity.”
Managed Compute deployment lifecycle:
  • “Use model_catalog_list to find models advertised for Managed Compute, where isManagedCompute is true.”
  • “For openai--gpt-oss-20b, use model_details_get to confirm that isFoundryManagedCompute is true, and show its registry modelAssetId, resolved deployment templates, and accelerators.”
  • “Use model_quota_list to check the selected accelerator in East US 2 before deployment.”
  • “Review the deployment template, accelerator, and instance count with me. After I confirm the billable dedicated GPU allocation, use model_deploy_managed_compute to create gpt-oss-managed-test with the registry modelAssetId, one resolved deploymentTemplate, and the approved instanceCount.”
  • “Use model_deployment_get to monitor gpt-oss-managed-test and verify that its Kind is Managed.”
  • “When testing is complete and I confirm cleanup, use model_deployment_delete to delete gpt-oss-managed-test.”
Agent evaluation workflow:
  • “List all agents in my project.”
  • “Evaluate my customer-support-agent v2 using Relevance, Groundedness, and Safety evaluators.”
  • “Compare my baseline evaluation against the new run.”
  • “Show me the comparison results with statistical significance.”
Continuous evaluation setup:
  • “List all agents in my project.”
  • “Enable continuous evaluation for my customer-support-agent using Relevance, Groundedness, and Safety evaluators with the existing gpt-5.6-luna deployment.”
  • “Show me the continuous evaluation configuration for customer-support-agent.”
  • “Update continuous evaluation to sample 25% of responses with a maximum of 10 runs per hour.”
Resource management and cleanup:
  • “List all my current deployments and their usage.”
  • “Check which deployments are using deprecated model versions.”
  • “Show me my quota usage across all regions.”
  • “Delete unused test deployments to free up capacity.”

Preview limitations

Foundry MCP Server is in public preview. The following limitations apply:
  • No network isolation — Foundry MCP Server uses the public endpoint https://mcp.ai.azure.com. Resources behind Azure Private Links aren’t accessible. For private MCP connectivity, build your own MCP server and connect it to Agent Service with private networking.
  • Data residency — Requests and responses might be processed in EU or US data centers. The server itself doesn’t store data, but cross-region processing can occur.
  • No SLA — Preview features don’t include a service-level agreement. Don’t use the server for production workloads that require guaranteed availability.
  • Tool set might change — Tools, parameters, and return values might change during the preview period without notice.
For more information, see Supplemental Terms of Use for Microsoft Azure Previews.

Common errors

For more troubleshooting guidance, see Foundry MCP Server security and best practices.