Models before the GPT-5.6 family don’t charge extra to write to the cache. On GPT-5.6 models and later model families, cache writes can incur charges in addition to discounted cache reads. To keep costs predictable, structure your prompts so that reused content stays identical across requests, which favors cache reads over cache writes. For current rates, see the Azure OpenAI pricing page.
Improve cache hit rates with a prompt cache key
On GPT-5.6 models and later model families, set theprompt_cache_key parameter and reuse the same key for requests that share long, common prompt prefixes. This parameter improves cache matching for related requests. You don’t need a specific API version to use prompt_cache_key. For new integrations, use the v1 API.
If requests for the same prefix and prompt_cache_key combination exceed approximately 15 requests per minute, some requests might miss the cache. For higher-volume workloads, distribute requests across multiple keys while keeping a stable mapping between each key and its shared prompt prefixes.
Configure prompt cache breakpoints
On GPT-5.6 models and later model families, use explicit cache breakpoints to mark the end of a reusable prompt prefix. Both the Responses API and Chat Completions API support breakpoints. Azure OpenAI uses the same request structures as the OpenAI APIs, but setmodel to your Azure model deployment name. Content after the breakpoint can change without invalidating the cached prefix.
Standard pay-as-you-go deployments support prompt cache breakpoints. Provisioned Throughput managed (PTU-M) deployments don’t support prompt cache breakpoints.
Set the request-wide cache policy by using prompt_cache_options.mode:
Set
prompt_cache_options.ttl to 30m to configure the minimum cache lifetime for all breakpoints in the request. The 30m value is the default and the only supported value. This setting doesn’t select the in-memory or extended retention policy.
Add prompt_cache_breakpoint: { "mode": "explicit" } to a supported prompt content block. The breakpoint includes the block and all prompt content before it in the reusable prefix.
- The Responses API supports breakpoints on
input_text,input_image, andinput_fileblocks. - The Chat Completions API supports breakpoints on
text,image_url,input_audio, andfileblocks.
Breakpoint limits
- Each request can create up to four new cache writes.
- In
implicitmode, the breakpoint on the latest message uses one write slot, so the request can write up to the latest three explicit breakpoints. - In
explicitmode, the request can write up to the latest four explicit breakpoints. - Breakpoints from earlier conversation turns are read-only. They can match the cache, but the request doesn’t write them again.
- For cache reads, Azure OpenAI considers up to the latest 50 breakpoints in the conversation.
implicit mode and adds an explicit breakpoint after a stable reference file:
explicit mode and marks the end of a reusable system message:
Models before the GPT-5.6 family don’t support
prompt_cache_options or prompt_cache_breakpoint. Requests that include these parameters return a 400 error. Continue to use automatic prompt caching with these models.Prompt cache retention
Prompt caching has two controls with different semantics:- On GPT-5.6 models and later model families,
prompt_cache_options.ttlsets a minimum cache lifetime. It doesn’t select a storage policy or maximum retention period. - For earlier models,
prompt_cache_retentionselects a maximum-retention policy. On GPT-5.6 models and later model families, this field doesn’t apply and is deprecated.
prompt_cache_options.ttl to set the minimum lifetime of all breakpoints written by the request. The only supported value is 30m, which is also the default. A cached prefix remains eligible for reuse for at least 30 minutes, but the service might retain it longer.
For models before the GPT-5.6 family, set prompt_cache_retention on your Responses or Chat Completions request. Prompt caching can use either in-memory or extended retention policies. When available, extended prompt caching aims to retain the cache for longer, so that subsequent requests are more likely to match the cache. Prompt cache pricing is the same for both policies.
In-memory prompt cache retention
The system typically clears caches within 5 to 10 minutes of inactivity and always removes them within one hour of the cache’s last use. The system doesn’t share prompt caches between Azure subscriptions. All Azure OpenAI models GPT-4o or newer support in-memory prompt cache retention. It applies to models that have chat-completion, completion, responses, or real-time operations. For models that don’t have these operations, this feature isn’t available.Extended prompt cache retention
Extended prompt cache retention keeps cached prefixes active for longer, up to a maximum of 24 hours. Extended prompt caching works by offloading the key/value tensors to GPU-local storage when memory is full, which significantly increases the storage capacity available for caching. Extended prompt cache retention is available for the following models:gpt-5.5gpt-5.4gpt-5.3-codexgpt-5.2gpt-5.1-codex-maxgpt-5.1gpt-5.1-codexgpt-5.1-codex-minigpt-5.1-chatgpt-5gpt-5-codexgpt-4.1
Configure per request
Forgpt-5.4 and older models, if you don’t specify a retention policy, the default is in_memory. Allowed values are in_memory and 24h. For gpt-5.5, extended retention is enabled by default.
Getting started
To take advantage of prompt caching, a request must meet both of these requirements:- A minimum of 1,024 tokens in length.
- The first 1,024 tokens in the prompt must be identical.
cached_tokens under prompt_tokens_details in the chat completions response.
On GPT-5.6 models and later model families, Standard pay-as-you-go deployments report cache reads in cached_tokens and cache writes in cache_write_tokens. The following excerpt shows these fields in a Chat Completions response. JSON property order isn’t significant and might vary.
cached_tokens value of 0. Prompt caching is enabled by default for supported models.
Best practices
- Place stable or repeated content at the beginning of the prompt and dynamic content at the end. Keep conversation context append-only.
- Reuse a consistent
prompt_cache_keyfor requests that share a prefix. For high-volume workloads, partition traffic across keys while keeping a stable mapping between each key and its prefixes. - On Standard pay-as-you-go deployments with GPT-5.6 models and later model families, place explicit breakpoints after stable content. Use
explicitmode when you want only the breakpoints you provide to be eligible for cache reads and writes. - Monitor cache reads with
cached_tokens. On Standard pay-as-you-go deployments with GPT-5.6 models and later model families, also monitor cache writes withcache_write_tokensand compare write volume with later cache reads. - Maintain a steady stream of requests with identical prefixes to improve cache reuse.
Frequently asked questions
The following answers clarify supported cache content, costs, deployment types, and data residency.What is cached?
Feature support for o1-series models varies by model. For more information, see the dedicated reasoning models guide. Prompt caching supports:
To improve the likelihood of cache hits, structure your requests so that repetitive content occurs at the beginning of the messages array.
Can I disable prompt caching?
On Standard pay-as-you-go deployments with GPT-5.6 models and later model families, setprompt_cache_options.mode to explicit and don’t add any explicit breakpoints. The request doesn’t use prompt caching or incur cache-write charges. Earlier models and PTU-M deployments don’t support this option; prompt caching remains enabled by default.
Do I pay extra to write to the cache?
On models before the GPT-5.6 family, there’s no extra charge to write to the cache. On GPT-5.6 models and later model families, cache writes can incur charges in addition to discounted cache reads. To see current rates, go to the Azure OpenAI pricing page.Do prompt cache breakpoints work with PTU-M?
On GPT-5.6 models and later model families, Standard pay-as-you-go deployments support prompt cache breakpoints and exposecache_write_tokens. Provisioned Throughput managed (PTU-M) deployments continue to support prompt caching, but they don’t support prompt cache breakpoints or expose cache_write_tokens.