Prerequisites
- Familiarity with the concepts in What is provisioned throughput for Foundry Models?
- An estimate of your workload characteristics: expected peak requests per minute (RPM), average prompt size in tokens, and average response size in tokens.
Estimate PTUs for text-only model
For a text model, the workload always includes text output tokens. Input can consist of text input tokens, image input tokens, or a combination of both. This section covers how to estimate the number of PTUs required for a workload when you use a text model. For how to estimate the number of PTUs for an image model, see Estimate PTUs for image model. Two approaches are available for estimating the PTUs required for a workload:- Use the sizing formulas for full control over the calculation
- Use the Foundry capacity calculator for a guided estimate.
For older models (before GPT-4o), the request/call shape distribution affects capacity consumption: a small number of large calls can consume significantly more capacity than many small calls with the same average token count. For GPT-4o and later models, TPM per PTU is set for input and output tokens separately, so this tiering effect doesn’t apply.
Estimate manually
You can estimate the PTUs your workload requires by using the model-specific values from the deployment parameters tables and information about your expected traffic as follows:Normalized TPM
The manual calculation of PTUs converts your expected token volume into a single number called the normalized TPM. The number of PTUs required is then determined by dividing the normalized TPM by the model’s Input TPM per PTU value. Formulas:- Input TPM = Peak RPM × average prompt size
- Output TPM = Peak RPM × average response size
- Normalized TPM = (input TPM × (1 − cache rate)) + (output-to-input ratio × output TPM)
- PTUs required = normalized TPM ÷ Input TPM per PTU
- Input TPM = 1,000 × 200 = 200,000
- Output TPM = 1,000 × 20 = 20,000
- Normalized TPM (no cache) = 200,000 + (8 × 20,000) = 360,000
- PTUs required = 360,000 ÷ 3,400 = 105.88 (110 PTUs rounded up to the nearest 5 PTUs, matching the Data Zone Provisioned scale increment for gpt-5.2.)
- Effective input TPM = 200,000 × (1 − 0.50) = 100,000
- Normalized TPM = 100,000 + (8 × 20,000) = 260,000
- PTUs required = 260,000 ÷ 3,400 = 76.47 (80 PTUs rounded up to the nearest 5 PTUs, matching the Data Zone Provisioned scale increment for gpt-5.2.)
1 Rounded up to the nearest 5 PTUs, matching the Data Zone Provisioned scale increment for gpt-5.2.
Use the capacity calculator
Use the capacity calculator in the Foundry portal to size specific workload shapes. Find the calculator on the Quota page and enter the following parameters based on your workload:
After you fill in the required details, select Calculate. The output shows:
- The estimated PTU count required for the workload. This value is rounded up to the nearest PTU scale increment for the selected deployment type, or to the deployment type’s minimum PTU count, depending on which one is larger.
- The raw (unrounded) estimated PTU count.
How input and output tokens affect throughput
The throughput that a deployment gets per PTU depends on the model and the mix of input and output tokens in a given minute. Throughput is measured as tokens per minute, or TPM. Generating output tokens requires more processing capacity than consuming input tokens. For GPT-4.1 models and later, the system determines an output-to-input ratio to match the global standard price ratio between input and output tokens, with exceptions for some models. For example,- For gpt-5, one output token counts as eight input tokens toward your utilization limit, matching the model’s global standard price ratio.
- For gpt-4.1, one output token counts as four input tokens.
- Older models use different ratios.
Models with a non-standard output-to-input ratio
Some models use an output-to-input ratio that differs from their global standard price ratio. For example, with Llama-3.3-70B-Instruct, one output token counts as four input tokens toward your utilization limit, which differs from that model’s standard price ratio. See pricing for Llama models for the full input and output pricing breakdown.Deployment parameters and throughput values by model
The tables in this section list the throughput and deployment parameters for each supported model. To understand what the parameters in each row mean, see Understand deployment parameters.Latest Azure OpenAI models
Long-context requests that exceed 128K prompt tokens aren’t supported for
gpt-5.4, gpt-4.1, gpt-4.1-mini, and gpt-4.1-nano. The system routes these requests to spillover deployments, if available. Otherwise, the requests return an error.
1 Calculated as p50 request latency on a per 5-minute basis. TPS = tokens per second.
2 Latency SLA not defined.
Normalized token pricing for GPT-6 Astra and newer models
GPT-6 Astra, GPT-6 Sol, GPT-6.1 Sol, and newer Azure OpenAI models use normalized tokens to align PTU capacity consumption with Global Standard pay-as-you-go token pricing. This method accounts for input, cached input, cache writes, and output, including different weights for short- and long-context requests. GPT-6 Astra and GPT-6 Sol support prompt cache breakpoints for explicit control over prompt caching. Cache reads and writes are included in normalized-token accounting.One short-context input token is one normalized token. Typically, calculate other token weights by dividing their prices by the same model’s short-context input price. For GPT-6.1 Sol, the current short-context cached-input weight differs from this price ratio, as noted in the table below.Cached input and cache writes consume PTU capacity at their respective weights. The older sizing rule that deducts cached input entirely doesn’t apply to models using this accounting method.
Pay-as-you-go prices and token weights
The following prices are in USD per 1 million tokens. GPT-6 Astra, GPT-6 Sol, and GPT-6.1 Sol have the same normalized weights except for GPT-6.1 Sol’s long-context cached input. Its short-context cached input also has a temporary PTU weight that differs from its price ratio.
1 The GPT-6.1 Sol short-context cached-input price corresponds to a weight of 0.05. Due to a current system limitation, PTU capacity uses a weight of 0.1 instead. Use 0.1 when sizing a deployment.
Normalized token cost = token-type price ÷ the model's short-context input price
For GPT-6 Astra, divide by 2.00. Use the same pricing unit in the numerator and denominator. For example, GPT-6 Sol long-context output has a weight of $15.00 ÷ $2.00 = 7.5. The GPT-6.1 Sol short-context cached-input weight is the exception described above.
Convert normalized tokens to PTUs
Each PTU supplies a model-specific normalized-token budget per minute. Because short-context input has a weight of 1, that budget equals the model’s Input TPM per PTU value.- Normalized TPM = Sum of (tokens per minute in each token category × that category’s normalized token cost).
- Raw PTUs required = Normalized TPM ÷ the model’s normalized TPM per PTU.
PTU and pay-as-you-go comparison
This illustration uses one PTU at $260.00 per month, retaining the existing article’s example cost and applying it to GPT-6 Astra and GPT-6 Sol. It assumes 100% sustained utilization for 30 days, no discounts, and the token prices above. It’s a capacity comparison, not a quote for a deployable configuration; deployment minimums still apply. GPT-6.1 Sol isn’t included because its current short-context cached-input PTU weight differs from its price ratio.
Under these assumptions, PTU and pay-as-you-go costs are approximately equal at full utilization. Using the same token weights makes this a token-mix-independent comparison of normalized-token capacity, not a guarantee of realized workload throughput or cost. Lower utilization increases the effective PTU cost per token.
For current Azure prices and purchase terms, see Azure OpenAI pricing.
Previous Azure OpenAI models
1 Calculated as the average request latency on a per-minute basis across the month. TPS = tokens per second.
Foundry Models sold by Azure
This section lists other Foundry models sold by Azure, not including the Azure OpenAI in Foundry Models listed in the previous tables.
1 For Llama-3.3-70B-Instruct, one output token counts as four input tokens toward your utilization limit. This ratio differs from the global standard price ratio between input and output tokens. See Models with a non-standard output-to-input ratio and Llama model pricing.
2 Calculated as the average request latency on a per-minute basis across the month. TPS = tokens per second.
Fireworks on Microsoft Foundry models
The following Fireworks on Microsoft Foundry models support both Global and US Data Zone provisioned throughput.
1 Calculated as the average request latency on a per-minute basis across the month. TPS = tokens per second.
Understand deployment parameters
Each row in the tables corresponds to one of the following parameters:Estimate PTUs for image model
Image models extend the PTU sizing methodology used with text-only models because a workload for an image model contains image output tokens in addition to text input tokens and image input tokens. Each of these token types consumes PTU capacity differently. To size a deployment, you must first convert all token types into a common unit called normalized tokens, which represent the equivalent number of text input tokens before estimating PTU requirements.Before using Azure Monitor to monitor a deployed image model, review Limitations and known issues for image model metrics.
Estimate PTUs manually
Estimate the PTUs your workload requires by using the throughput values for your image model and information about your expected traffic as follows:Traffic information
1 To estimate image input tokens per request, refer to the OpenAI documentation: Image input cost calculator.
2 To estimate image output tokens per request, refer to the OpenAI documentation: Image output token calculator.
Throughput values for image model
GPT-image-2 uses the following throughput values:
1 The image-to-text conversion factor is
Text input tokens per minute per PTU ÷ Image input tokens per minute per PTU. This factor means that one image token uses the same PTU capacity as about 1.6 text input tokens.
2 Image output tokens have a higher weight than image input tokens because image generation needs extra processing. The image output weighting happens before converting to text-token equivalents.
Normalized TPM
The manual calculation converts image input and output tokens to equivalent text input tokens. The resulting normalized TPM determines the number of PTUs required.- Normalized tokens per request = Text input tokens + (image input tokens × 1.6) + (image output tokens × 3.75 × 1.6)
- Normalized TPM = Normalized tokens per request × Peak RPM
- PTUs required = Normalized TPM ÷ 1,200
- Normalized tokens per request = 2,000 + (1,229 × 1.6) + (7,024 × 3.75 × 1.6) = 46,110.4
- Normalized TPM = 46,110.4 × 10 = 461,104
- PTUs required = 461,104 ÷ 1,200 = 384.25 (400 PTUs rounded up to the nearest 100 PTUs, matching the scale increment for GPT-image-2.)
1 Rounded up to the 100-PTU scale increment for GPT-image-2.