model-router in the model field to let Foundry pick the best model automatically. Or pass a specific model name for deterministic control. The code is the same—only the model value changes.
You pass your deployment name to the
model parameter. In most cases the deployment name matches the model name. For example, a gpt-4.1-mini deployment is called "gpt-4.1-mini".Prerequisites
- A Foundry project with a
model-routerdeployment. See Deploy model router. - At least one named model deployment for deterministic calls (for example,
gpt-4.1-mini). See Deploy a model. - Familiarity with the Responses API.
- Python 3.10+ or Node.js 22+.
- The Foundry SDK for your language:
Call models through the Responses API
The following sample calls several models through the sameresponses.create() interface, starting with model-router for automatic selection, then named models for deterministic control.
- Reference:
AIProjectClient.get_openai_client(Python) /AIProjectClient.getOpenAIClient(JavaScript/TypeScript) - Reference: OpenAI Responses API (
responses.create, both languages)
See the first row:
model-router didn’t target a specific model, but the Responded column shows that it selected gpt-4.1-nano. For the named models that follow, the two columns match. The code is identical in every case.
Routing strategies
Every call goes throughresponses.create(). The model value is the only decision point.
Use
model-router as your default. Customize your model router deployment with optional settings. See Model router deployment options.
Switch to a named model only when you need deterministic control.
When you evaluate automatic routing, use the named model that currently serves your workload as the baseline. Keep the request and application configuration consistent, then compare the quality, cost, and latency outcomes before you adopt model router broadly. See Evaluate model router for your workload.
Built-in enterprise capabilities
Everyresponses.create() call, whether routed through model-router or targeting a named model, automatically includes:
- Automatic failover—When using
model-router, if the selected model encounters a transient issue, model router transparently redirects the request to the next most appropriate model. No disruption to your application, no retry logic required. If you configure a model subset, that subset also serves as your fallback set — select at least two models to benefit from failover. - Prompt caching—Model router supports prompt caching. When model router delegates a request to a model that supports prompt caching, cached tokens are used automatically. Combined with model router’s right-fit model selection, you get an extra efficiency lift: the optimal model for the task and reduced token costs on repeated prompt prefixes — no configuration needed.
- Content filtering—Configurable content safety applied to inputs and outputs without extra API parameters.
- Role-based access control—Azure role-based access control governs who can call which deployments. No separate API key management.
- Model identification—The response includes
response.model, which identifies the model that handled the request. To inspect preview routing attempts, status, and reported latency for Chat Completions requests, see Monitor model router. - Data residency and compliance—Traffic stays within your Azure region. No data leaves your tenant boundary.
- Rate limiting and quotas—Per-deployment token-per-minute limits protect your workloads from noisy neighbors.
Related content
- Use model router for Microsoft Foundry — deployment, routing modes, model subsets
- Use the Azure OpenAI Responses API — streaming, tools, stored responses
- Microsoft Foundry SDKs and endpoints — client setup and endpoint patterns