Deployment options
Foundry provides two deployment options:- Serverless API — For Foundry Models, including Foundry Models sold by Azure and select Models from partners and community. This option is the preferred and most capable deployment path. It includes the standard, provisioned throughput, batch, and developer deployment types.
- Managed compute (preview) — For open-source, partner, and custom models that run on dedicated GPU capacity that Foundry manages for you.

Serverless API
Serverless API is the preferred deployment option in Foundry. It supports the widest range of capabilities and deployment types.Which models use serverless API deployments?
All Foundry Models, including Foundry Models sold by Azure and select Models from partners and community, use serverless API deployments. Foundry Models sold by Azure include all Azure OpenAI models and selected models from top providers that are billed through your Azure subscription, covered by Azure service-level agreements, and supported by Microsoft. Models from partners and community that use serverless API deployments include Anthropic models and specific models from partners like Mistral, Cohere, and Meta.Serverless API capabilities
Serverless API deployments support:- Multiple deployment types (or deployment SKUs) — Global Standard, Data Zone Standard, Standard (single region), provisioned, batch, and more. Each type controls where data is processed and how you pay. For details, see Deployment types for Microsoft Foundry Models.
- Data processing flexibility — Choose regional, data zone (US, EU, or APAC), or global processing based on your compliance requirements.
- Content filtering — Built-in Azure AI Content Safety filters with customizable configurations.
- Keyless authentication — Microsoft Entra ID (recommended) and key-based authentication.
- Private networking — Virtual network integration for secure access.
- Provisioned throughput — Reserve capacity with provisioned throughput units (PTUs) for predictable, low-latency performance. For details, see Provisioned throughput.
Resource requirements
Serverless API deployments are available in:- Foundry resources — The primary resource type for new Foundry projects. No AI Hub required.
- Azure OpenAI resources — If you use Azure OpenAI resources, the model catalog shows only Azure OpenAI models for deployment. Upgrade to a Foundry resource for access to the full set of Foundry Models.
Managed compute deployment (preview)
Managed compute in Foundry is currently in public preview.
This preview is provided without a service-level agreement, and we don’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
Managed compute supports open-source, partner, industry, and custom models. Managed compute deployments are served on the unified Foundry project endpoint, using the same authentication, networking, and SDK surface.
Which models use managed compute?
You can deploy models from the Hugging Face Collection by using managed compute. Examples include:- Qwen models
- NVIDIA Nemotron models
- Selected Meta models
- Selected Mistral models
Managed compute capabilities
Managed compute (Preview) supports:- Unified Foundry endpoint and authentication — Use the same project endpoint, API keys, Microsoft Entra ID, and private networking as pay-per-token and provisioned throughput deployments. Inference routes use
<endpoint>/managed-deployments/<deployment-name>/. Chat-completions-compatible runtimes also work on the standard/openai/v1/route with the OpenAI SDK. - Model-instance sizing — Deployments are sized in model-centric terms. You don’t need to pick virtual machine SKUs, because Foundry chooses GPUs per instance based on model size, architecture, context length, and whether the workload is optimized for latency or throughput.
- Optimized inference runtimes — Microsoft-curated vLLM, SGLang, and NVIDIA NIM containers with continuous batching and tensor parallelism.
- Accelerator families — A100 (80 GB), H100 (80 GB), and MI300X (192 GB).
- Microsoft-managed runtimes — Microsoft owns serving runtimes, base container images, and security patches. Updates are applied to live deployments automatically.
- Observability metrics — Each deployment emits API call count by status code and response-time percentiles. Chat-completion models also emit input and output token counts, time-to-first-token (TTFT) percentiles, and total response-time percentiles, grouped by time.
Billing and quota
Managed compute billing is hourly per accelerator SKU, with throughput per GPU as the underlying billing unit. Quota is granted per accelerator SKU per region through the Foundry quota process and is separate from Azure VM quota. Azure virtual machines are an infrastructure-as-a-service (IaaS) offering with regional SKUs; managed compute is a PaaS offering that leads with Global and Data Zone processing. Existing Azure VM quota can’t be applied to a managed compute deployment. Managed compute is currently available for global deployment. For rate estimates, see the Azure pricing calculator.Get started
To get started with managed compute deployment, see Deploy open-source models with managed compute.Deployment option comparison
Use Serverless API whenever possible. The following table compares capabilities across the two deployment options:Instant access (preview) isn’t a deployment option in this comparison. It calls supported models by name without creating a Serverless API or managed compute deployment.