Claude quota scope
Microsoft Foundry manages Claude model quota at the subscription level. Resources and regions share quota instead of receiving separate allocations.- All Global Standard deployments of the same model and version in a subscription draw from one shared quota pool across all regions.
- All Data Zone Standard deployments of the same model and version in a subscription draw from a shared quota pool within each data zone, such as the US data zone.
Rate-limit measurements
Claude models use the following rate-limit measurements for each model:- Requests per minute (RPM) measures the number of requests.
- Uncached input tokens per minute (ITPM) measures input tokens that aren’t read from a prompt cache.
- Output tokens per minute (OTPM) measures tokens that the model generates.
Cache-aware ITPM
For most Claude models, only uncached input tokens count toward ITPM limits. These tokens include:- Input tokens: Tokens in the request after the last cache breakpoint (uncached input).
- Cache creation input tokens: Tokens written to either the 5-minute or 1-hour prompt cache.
Default rate limits by subscription type
Your Azure subscription type determines your default rate limits. The Version 2: Hosted on Azure and Version 1: Hosted on Anthropic infrastructure columns indicate whether each model and deployment type combination supports quota allocation. Yes means the combination supports quota allocation, but its default numeric limit might be zero. N/A means the combination doesn’t support quota allocation for that deployment type. The RPM, ITPM, and OTPM columns show the default capacity.- Pay-as-you-go
- Enterprise and MCA-E
- Free Trial