service_tier to flex.
Use Flex processing for noninteractive and lower-priority work, such as model evaluations, data enrichment, document analysis, and asynchronous application workflows. For latency-sensitive or capacity-sensitive workloads, use Standard processing, Priority processing, or provisioned throughput instead.
With the introduction of Flex processing, requests that set
service_tier to flex are processed only when the selected model supports Flex processing. An unsupported model returns an HTTP 400 invalid_request_error and doesn’t fall back to Standard processing. Flex processing has no latency SLA or service SLA.Prerequisites
- An Azure subscription. Create one for free.
- An Azure OpenAI resource with a supported model deployed by using the Global Standard deployment type.
-
The resource endpoint and an API key or Microsoft Entra ID credentials. The examples in this article use an API key stored in the
AZURE_OPENAI_API_KEYenvironment variable. - A workload that can tolerate variable latency and transient resource-unavailable responses.
-
Python 3.10 or later and the OpenAI Python package for the Python examples:
Send a Flex request
Setservice_tier to flex in each request that should use Flex processing. The model value is the name of your Azure model deployment.
Python
The following example sends a Flex request by using the Responses API:REST
The following example sends the same request directly to the Responses API:service_tier field in a successful response.
Reference: Responses API REST reference
Choose a processing option
Flex, Standard, and Priority processing are service-tier choices for online API requests. Batch and provisioned throughput are separate deployment and purchasing options.
Choose Flex processing when all of the following conditions apply:
- Your workload can tolerate longer and variable processing times.
- You prefer lower cost over predictable latency.
- Your application can retry transient failures or route a failed request to Standard processing.
- The selected model and request context are supported.
- A user is waiting for an interactive response.
- The request must complete within a strict latency target.
- Your application can’t tolerate or retry transient HTTP 429 responses.
- You require reserved processing capacity or predictable throughput.
- Request-level selection: Set
service_tiertoflexon each request that should use Flex processing. - No separate deployment: Send Standard and Flex requests to the same Global Standard deployment, and select the tier per request.
- Supported APIs: Use the Responses API or Chat Completions API.
- Synchronous response: The API call remains synchronous, even though the workload can take longer to complete. Flex processing isn’t the same as the Batch API.
- Capacity-dependent availability: A request can return HTTP 429 when Flex capacity isn’t available.
- No automatic Standard fallback: Your application must explicitly retry with
service_tierset todefaultif Standard processing is acceptable. - Shared quota: Flex and Standard requests use the quota assigned to the Global Standard deployment.
- Same model output quality: Flex uses the same underlying model as Standard. The processing tier changes latency, availability, and price, not model quality.
Flex input and output tokens receive a 50% discount compared with the corresponding Standard token rates. Eligible cached input tokens also receive the applicable cached-token discount. For current rates, see Azure OpenAI pricing.
Review supported models
Flex processing has limited model availability at launch.gpt-5.6-sol is the first supported model. The following table lists supported models. Microsoft adds more models as support becomes available.
Check this table before you send a Flex request. Don’t assume that a model or a new model version supports Flex processing because it supports Standard or Priority processing. An unsupported model returns HTTP 400. To avoid disrupting your application, implement an application-level fallback to Standard processing when Standard pricing and performance are acceptable.
Fall back to Standard processing
Flex processing doesn’t automatically route a request to Standard when Flex capacity is unavailable. If completing the request is more important than retaining Flex pricing, retry the request withservice_tier set to default.
The following example uses a completion-first policy. It first attempts Flex processing and retries once with Standard processing after any HTTP 429 response:
RateLimitError in your application. When the service returns the Flex-specific error identifier, use it to limit fallback to capacity-related responses.
For workloads that prioritize the lowest cost, retry Flex processing with exponential backoff before falling back. For workloads that prioritize completion time, fall back to Standard after the first Flex capacity error.
Reference: RateLimitError
Handle Flex errors
Distinguish permanent request errors from transient capacity errors.A Flex request rejected because processing capacity is unavailable isn’t billed. However, you might notice less available rate-limit capacity because Flex and Standard requests share the quota assigned to the Global Standard deployment.
- Honor
Retry-Afterwhen the response includes it. - Otherwise, wait for an exponentially increasing delay with random jitter.
- Retry Flex only while the delay remains acceptable for the workload.
- Fall back to Standard if the retry budget is exhausted and the application allows the higher Standard cost.
- Return an explicit failure if neither delayed Flex processing nor Standard fallback meets the application’s requirements.
Monitor usage and costs
Use Azure Monitor metrics to compare Flex and Standard traffic on the same deployment. Monitor request volume, token consumption, latency, failures, and the rate at which Flex requests fall back to Standard in your application.- Sign in to the Azure portal.
- Go to your Azure OpenAI resource, and select Metrics.
- Add the Azure OpenAI Requests metric. You can also add Azure OpenAI Latency, Azure OpenAI Usage, and error metrics.
- Add a filter where ServiceTierRequest equals
flex.

- Create alerts for sustained HTTP 429 responses, increased error rates, and latency that exceeds your workload’s retry budget.
For more information about monitoring model deployments, see Monitor Azure OpenAI.
Flex usage is billed on dedicated Flex meters so that you can distinguish it from Standard usage. Use Cost Analysis to review Flex token costs by resource and deployment.
- In the Azure portal, open Cost Management + Billing > Cost analysis.
- Filter to the subscription, resource group, or Azure OpenAI resource that contains the deployment.
- Group or filter by Meter to separate Flex usage from Standard usage.
- Add a billing Tag filter, select deployment, and choose the deployment name.
- Compare Flex cost savings with Standard fallback costs and the workload’s completion requirements.
Apply production best practices
- Set a longer timeout. Flex requests can take longer than Standard requests. Start with a client timeout appropriate for your workload, such as 15 minutes, and test with representative prompts.
- Use bounded retries. Limit retry attempts and total elapsed time.
- Add jitter. Randomize backoff delays to avoid synchronized retry spikes.
- Make fallback explicit. Set
service_tiertodefaultrather than relying on implicit behavior. - Track the processed tier. Record the response
service_tiervalue with latency, status, token usage, and cost data. - Separate interactive and background traffic. Keep user-facing requests on Standard, Priority, or provisioned throughput unless variable Flex latency is acceptable.
- Control duplicate work. Ensure the application doesn’t submit the same logical job multiple times after client-side timeouts.
- Test failure paths. Validate handling for HTTP 400, 408, 429, and transient 5xx responses before using Flex processing in production workflows.
- Review model support before upgrades. A replacement model or model version doesn’t automatically inherit Flex support.
Override the service tier with a request header
Use thex-ms-service-tier request header when a gateway, proxy, or centralized
routing layer needs to select the service tier without inspecting or modifying
the request body. The header can also reduce migration changes for applications
that already select an OpenAI service tier through a request header.
The header accepts the following values:
When the header is present and valid, it takes precedence over the
service_tier value in the request body.
This example requests Flex processing through the header. The header overrides
the
default value in the request body: