Skip to main content
Flex processing (preview) provides inference at a 50% discount compared with Standard processing for workloads that can tolerate slower response times and occasional resource unavailability. Select Flex processing for an individual Responses API or Chat Completions API request by setting service_tier to flex. Use Flex processing for noninteractive and lower-priority work, such as model evaluations, data enrichment, document analysis, and asynchronous application workflows. For latency-sensitive or capacity-sensitive workloads, use Standard processing, Priority processing, or provisioned throughput instead.
With the introduction of Flex processing, requests that set service_tier to flex are processed only when the selected model supports Flex processing. An unsupported model returns an HTTP 400 invalid_request_error and doesn’t fall back to Standard processing. Flex processing has no latency SLA or service SLA.

Prerequisites

  • An Azure subscription. Create one for free.
  • An Azure OpenAI resource with a supported model deployed by using the Global Standard deployment type.
  • The resource endpoint and an API key or Microsoft Entra ID credentials. The examples in this article use an API key stored in the AZURE_OPENAI_API_KEY environment variable.
  • A workload that can tolerate variable latency and transient resource-unavailable responses.
  • Python 3.10 or later and the OpenAI Python package for the Python examples:

Send a Flex request

Set service_tier to flex in each request that should use Flex processing. The model value is the name of your Azure model deployment.

Python

The following example sends a Flex request by using the Responses API:
The response contains the generated analysis and the service tier that processes the request. Reference: Responses API

REST

The following example sends the same request directly to the Responses API:
To verify which tier processes the request, check the service_tier field in a successful response. Reference: Responses API REST reference

Choose a processing option

Flex, Standard, and Priority processing are service-tier choices for online API requests. Batch and provisioned throughput are separate deployment and purchasing options. Choose Flex processing when all of the following conditions apply:
  • Your workload can tolerate longer and variable processing times.
  • You prefer lower cost over predictable latency.
  • Your application can retry transient failures or route a failed request to Standard processing.
  • The selected model and request context are supported.
Don’t use Flex processing when any of the following conditions apply:
  • A user is waiting for an interactive response.
  • The request must complete within a strict latency target.
  • Your application can’t tolerate or retry transient HTTP 429 responses.
  • You require reserved processing capacity or predictable throughput.
Flex processing has the following characteristics:
  • Request-level selection: Set service_tier to flex on each request that should use Flex processing.
  • No separate deployment: Send Standard and Flex requests to the same Global Standard deployment, and select the tier per request.
  • Supported APIs: Use the Responses API or Chat Completions API.
  • Synchronous response: The API call remains synchronous, even though the workload can take longer to complete. Flex processing isn’t the same as the Batch API.
  • Capacity-dependent availability: A request can return HTTP 429 when Flex capacity isn’t available.
  • No automatic Standard fallback: Your application must explicitly retry with service_tier set to default if Standard processing is acceptable.
  • Shared quota: Flex and Standard requests use the quota assigned to the Global Standard deployment.
  • Same model output quality: Flex uses the same underlying model as Standard. The processing tier changes latency, availability, and price, not model quality.
Flex input and output tokens receive a 50% discount compared with the corresponding Standard token rates. Eligible cached input tokens also receive the applicable cached-token discount. For current rates, see Azure OpenAI pricing.

Review supported models

Flex processing has limited model availability at launch. gpt-5.6-sol is the first supported model. The following table lists supported models. Microsoft adds more models as support becomes available. Check this table before you send a Flex request. Don’t assume that a model or a new model version supports Flex processing because it supports Standard or Priority processing. An unsupported model returns HTTP 400. To avoid disrupting your application, implement an application-level fallback to Standard processing when Standard pricing and performance are acceptable.

Fall back to Standard processing

Flex processing doesn’t automatically route a request to Standard when Flex capacity is unavailable. If completing the request is more important than retaining Flex pricing, retry the request with service_tier set to default. The following example uses a completion-first policy. It first attempts Flex processing and retries once with Standard processing after any HTTP 429 response:
Fallback to Standard changes the request’s pricing and performance characteristics. Use this pattern only when the workload can accept Standard pricing. An HTTP 429 response can indicate unavailable Flex capacity or a quota limit. Because Flex and Standard processing share quota, the Standard request might also fail when quota caused the original response. Apply retry limits and handle a second RateLimitError in your application. When the service returns the Flex-specific error identifier, use it to limit fallback to capacity-related responses. For workloads that prioritize the lowest cost, retry Flex processing with exponential backoff before falling back. For workloads that prioritize completion time, fall back to Standard after the first Flex capacity error. Reference: RateLimitError

Handle Flex errors

Distinguish permanent request errors from transient capacity errors.
A Flex request rejected because processing capacity is unavailable isn’t billed. However, you might notice less available rate-limit capacity because Flex and Standard requests share the quota assigned to the Global Standard deployment.
Use exponential backoff with random jitter for HTTP 408, 429, and transient 5xx responses. Set a maximum retry count and maximum delay so that a failed request doesn’t remain in an unbounded retry loop.
  1. Honor Retry-After when the response includes it.
  2. Otherwise, wait for an exponentially increasing delay with random jitter.
  3. Retry Flex only while the delay remains acceptable for the workload.
  4. Fall back to Standard if the retry budget is exhausted and the application allows the higher Standard cost.
  5. Return an explicit failure if neither delayed Flex processing nor Standard fallback meets the application’s requirements.
Don’t repeatedly retry the same unsupported Flex request. A retry succeeds only after you change the model, context length, API configuration, or service tier.

Monitor usage and costs

Use Azure Monitor metrics to compare Flex and Standard traffic on the same deployment. Monitor request volume, token consumption, latency, failures, and the rate at which Flex requests fall back to Standard in your application.
  1. Sign in to the Azure portal.
  2. Go to your Azure OpenAI resource, and select Metrics.
  3. Add the Azure OpenAI Requests metric. You can also add Azure OpenAI Latency, Azure OpenAI Usage, and error metrics.
  4. Add a filter where ServiceTierRequest equals flex.
Screenshot of Azure Monitor metrics filtered to Flex requests by using the ServiceTierRequest property.
  1. Create alerts for sustained HTTP 429 responses, increased error rates, and latency that exceeds your workload’s retry budget.
Track the following signals for each workload: For more information about monitoring model deployments, see Monitor Azure OpenAI. Flex usage is billed on dedicated Flex meters so that you can distinguish it from Standard usage. Use Cost Analysis to review Flex token costs by resource and deployment.
  1. In the Azure portal, open Cost Management + Billing > Cost analysis.
  2. Filter to the subscription, resource group, or Azure OpenAI resource that contains the deployment.
  3. Group or filter by Meter to separate Flex usage from Standard usage.
  4. Add a billing Tag filter, select deployment, and choose the deployment name.
  5. Compare Flex cost savings with Standard fallback costs and the workload’s completion requirements.
Flex input and output tokens are priced at 50% of the corresponding Standard rates. Prompt caching can reduce the cost of eligible cached input tokens further. A Flex request rejected because processing capacity is unavailable isn’t billed.

Apply production best practices

  • Set a longer timeout. Flex requests can take longer than Standard requests. Start with a client timeout appropriate for your workload, such as 15 minutes, and test with representative prompts.
  • Use bounded retries. Limit retry attempts and total elapsed time.
  • Add jitter. Randomize backoff delays to avoid synchronized retry spikes.
  • Make fallback explicit. Set service_tier to default rather than relying on implicit behavior.
  • Track the processed tier. Record the response service_tier value with latency, status, token usage, and cost data.
  • Separate interactive and background traffic. Keep user-facing requests on Standard, Priority, or provisioned throughput unless variable Flex latency is acceptable.
  • Control duplicate work. Ensure the application doesn’t submit the same logical job multiple times after client-side timeouts.
  • Test failure paths. Validate handling for HTTP 400, 408, 429, and transient 5xx responses before using Flex processing in production workflows.
  • Review model support before upgrades. A replacement model or model version doesn’t automatically inherit Flex support.

Override the service tier with a request header

Use the x-ms-service-tier request header when a gateway, proxy, or centralized routing layer needs to select the service tier without inspecting or modifying the request body. The header can also reduce migration changes for applications that already select an OpenAI service tier through a request header. The header accepts the following values: When the header is present and valid, it takes precedence over the service_tier value in the request body. This example requests Flex processing through the header. The header overrides the default value in the request body:
The override header doesn’t provide automatic fallback. An invalid header value or a tier that the selected deployment doesn’t support returns HTTP 400.