- Complex Code Generation: Capable of generating algorithms and handling advanced coding tasks to support developers.
- Advanced Problem Solving: Ideal for comprehensive brainstorming sessions and addressing multifaceted challenges.
- Complex Document Comparison: Perfect for analyzing contracts, case files, or legal documents to identify subtle differences.
- Instruction Following and Workflow Management: Particularly effective for managing workflows requiring shorter contexts.
Prerequisites
- An Azure OpenAI reasoning model deployed.
-
If you use the REST examples:
- Install the Azure CLI. For more information, see Install the Azure CLI.
-
Sign in with
az login, then generate a bearer token and store it in theAZURE_OPENAI_AUTH_TOKENenvironment variable.
Usage
These models don’t currently support the same set of parameters as other models that use the chat completions API.Chat completions API
- C#
- Python
- REST
- Output
How reasoning works
Reasoning models generate reasoning tokens in addition to the input and output tokens you’re already familiar with. The model uses those tokens to work through your prompt: breaking the problem apart, weighing approaches, and abandoning paths that don’t hold up. Reasoning tokens never appear in the message content, but they occupy space in the context window and are billed as output tokens. To see how many reasoning tokens a request consumed, checkcompletion_tokens_details.reasoning_tokens in a Chat Completions API response, or output_tokens_details.reasoning_tokens in a Responses API response.
The gpt-5.4 and gpt-5.5 models support interleaved thinking with the Responses API. They can produce visible output before and between periods of reasoning, and reason between tool calls.
Across a multi-turn conversation, input and output tokens carry forward from each turn. What happens to the reasoning from earlier turns depends on the model and on the reasoning.context value you set.
Manage the context window
Reasoning tokens share the context window with your input and the visible output. A single request can spend anywhere from a few hundred to tens of thousands of reasoning tokens depending on how hard the problem is, so leave room for them when you size a request. The usage object reports the exact count for each request:Control costs
Reasoning tokens are billed as output tokens, so a request that thinks longer costs more even when the visible answer is short. To cap the total the model generates, setmax_output_tokens with the Responses API or max_completion_tokens with the Chat Completions API. Both limits cover reasoning tokens, visible output tokens, and formatting tokens.
Capping output addresses only half of a multi-turn workload. Reasoning models also resend a growing conversation on every turn, and all_turns adds earlier reasoning items on top of that. To reduce what you pay for those repeated input tokens, see Prompt caching.
Allocate space for reasoning
If generation reaches the context window limit or the token cap you set, the response comes back incomplete:status on every response so your application handles this case instead of treating it as an empty result.
To avoid running out of room, reserve at least 25,000 tokens for reasoning and output while you’re getting a feel for a workload. Once you know how many reasoning tokens your prompts typically consume, tune the buffer to match.
Keep reasoning items in context
When a reasoning model calls functions through the Responses API, pass the reasoning items from the previous response back along with your function output. If the model called several functions in a row, send every reasoning item, function call item, and function call output item since the last user message. The model then continues the same line of reasoning instead of starting over, which reaches a good answer in fewer tokens. The simplest approach is to pass all output items from the previous response into the next request, either withprevious_response_id or by copying the items into the next input array. Reasoning items that aren’t relevant to your functions are ignored, and the relevant ones are retained.
If you trim or reorder context before sending it, keep everything between the last user message and your function call output intact.
Reasoning effort
Thereasoning_effort parameter controls how much the model thinks before it answers. Supported values vary by model and include none, minimal, low, medium, high, xhigh, and max. Defaults vary by model as well. For the values each model accepts, see API and feature support.
Reasoning models adapt within a setting, spending fewer tokens on simple tasks and thinking harder on complex ones. The higher the effort, the longer the model spends on the request, which generally produces more reasoning tokens.
o1-mini doesn’t support reasoning_effort.Developer messages
Developer messages ("role": "developer") are functionally the same as system messages.
Adding a developer message to the previous code example would look as follows:
- C#
- Python
- REST
- Output
Tool calling with reasoning models
Use the Responses API when you combine reasoning with function or custom tools. Thegpt-5.6 models support the Chat Completions API and tools, but can’t combine reasoning with tools on Chat Completions. A Chat Completions request that includes tools fails with the following error:
reasoning_effort, because these models default to medium. Sending tools is enough to trigger the error. An application that calls tools through Chat Completions can start failing after you upgrade its deployment from an earlier reasoning model.
For gpt-5.6 models, you have two ways to resolve it:
- Recommended: Send tool-calling requests to the Responses API. This path supports the model’s full range of
reasoning_effortvalues and returns reasoning items you can carry across turns. For a migration walkthrough, see Upgrade your Azure OpenAI app from Chat Completions to the Responses API. - If you must stay on Chat Completions, set
reasoning_efforttononeon every request that sendstools. The model then calls tools without reasoning, which loses the planning quality that reasoning provides.
ChatCompletionOptions.ReasoningEffortLevel:
ChatReasoningEffortLevel is marked experimental in the OpenAI .NET library, so it emits the OPENAI001 diagnostic. Suppress it with #pragma warning disable OPENAI001 as shown in the earlier samples, or add <NoWarn>$(NoWarn);OPENAI001</NoWarn> to your project file. Token-based authentication uses the same diagnostic.Reasoning mode
Thegpt-5.6 models support two execution modes in the Responses API. Standard mode is the default on Azure OpenAI. Set reasoning.mode to pro for difficult tasks that justify more model work and can absorb the extra latency.
Mode and effort are independent controls. The mode selects standard or pro execution, and reasoning_effort controls how much reasoning the model applies within that mode.
Reasoning summary
When using the latest reasoning models with the Responses API you can use the reasoning summary parameter to receive summaries of the model’s chain of thought reasoning. Thereasoning.summary parameter isn’t supported when multi-agent orchestration is enabled.
Attempting to extract raw reasoning through methods other than the reasoning summary parameter are not supported, may violate the Acceptable Use Policy, and may result in throttling or suspension when detected.
- C#
- Python
- REST
- Output
Even when enabled, reasoning summaries are not guaranteed to be generated for every step/request. This is expected behavior.
Preserve reasoning across calls
Conversation state and reasoning state aren’t the same thing. Passing messages across calls gives the model the visible conversation history. Persisted reasoning goes a step further: on models that support it, the model can also render its own reasoning items from earlier turns into the current context. Persisted reasoning is about continuity, not transparency. The reasoning items stay opaque, and the API never returns their reasoning text. Setreasoning.context to control which of the available reasoning items the model can draw on.
The GPT-5.4, GPT-5.5, and GPT-5.6 models support
all_turns. GPT-5.6 models use it by default, while GPT-5.4 and GPT-5.5 models default to current_turn.
Because
all_turns renders more reasoning items into context, it increases the tokens billed for a request. If you upgrade an existing workload to a GPT-5.6 model, expect higher token consumption on multi-turn conversations even when your code doesn’t change. Set reasoning.context to current_turn to keep the earlier behavior.- Setting
reasoning.contextdoesn’t create reasoning items that aren’t already available. It only controls which existing items the model renders. all_turnshas an effect only when the request can reach earlier response items. Useprevious_response_id, attach the response to a conversation, or replay the complete response history yourself.- On the first request in a conversation,
current_turnandall_turnsbehave the same way, because no earlier reasoning exists yet. - Each response reports the mode it actually used in its
reasoning.contextfield, as eithercurrent_turnorall_turns. Check that field to confirm the effective mode.
Continue reasoning with stored responses
When you store responses,previous_response_id is the shortest way to make earlier reasoning available to the model.
- C#
- Python
- REST
- Output
A C# example for
reasoning.context isn’t available yet. Select the Python or REST tab to see how to set the mode and read the effective value back from the response.current_turn when you replay older response items that the model no longer needs. Those items can stay in the request payload for continuity, but the service doesn’t render them into the new sample, which reduces the rendered context in long-running workflows.
Preserve reasoning without stored responses
In stateless mode, reasoning items in the response’soutput array include an encrypted_content property by default. Stateless mode applies when you set store to false, and when your organization uses Zero Data Retention. You don’t need to request the property: the API still accepts reasoning.encrypted_content in the include parameter for compatibility, but no longer requires it.
To use all_turns in this mode, keep every output item, append the next user message, and replay the complete history.
- C#
- Python
- REST
- Output
A C# example for stateless persisted reasoning isn’t available yet. Select the Python or REST tab to see how to replay encrypted reasoning items across turns.
Phase parameter
In long-running or tool-heavy workflows that usegpt-5.5 and gpt-5.4 in the Responses API, mark each assistant message with a phase value. The parameter is optional, but omitting it can cause the model to treat a preamble as the final answer and stop early.
Use commentary for intermediate assistant updates, such as the preamble a model produces before a tool call, and final_answer for the completed response. Don’t add phase to user messages.
previous_response_id, the service preserves the earlier assistant state for you. If you replay assistant history yourself, keep each message’s original phase value.
Python lark
GPT-5 series reasoning models have the ability to call a newcustom_tool called lark_tool. This tool is based on Python lark and can be used for more flexible constraining of model output.
Responses API
Chat Completions
Availability
Region availability
API and feature support
Input and output limits share the available context budget and aren’t additive. For details and a GPT-5.5 calculation example, see Understand model token limits and Responses API token budget.- GPT-6 reasoning models
- GPT-5 reasoning models
- O-Series Reasoning Models
- To avoid timeouts background mode is recommended for
o3-pro. o3-prodoes not currently support image generation.
Unsupported parameters
Reasoning models other than GPT-6 Astra don’t support the following parameters:temperature,top_p,presence_penalty,frequency_penalty,logprobs,top_logprobs,logit_bias,max_tokens
Prompting guidance
Reasoning models work best when you give them a clear goal, firm constraints, and an explicit output contract. Unlike non-reasoning models, they don’t need you to prescribe every intermediate step.- State the task, the constraints, and the output format you expect.
- Treat
reasoning_effortas a tuning knob rather than the first thing you reach for when quality drops. - For agentic or research-heavy workflows, define what counts as done and how the model should verify its own work.
Markdown output
By default theo3-mini and o1 models will not attempt to produce output that includes markdown formatting. A common use case where this behavior is undesirable is when you want the model to output code contained within a markdown code block. When the model generates output without markdown formatting you lose features like syntax highlighting, and copyable code blocks in interactive playground experiences. To override this new default behavior and encourage markdown inclusion in model responses, add the string Formatting re-enabled to the beginning of your developer message.
Adding Formatting re-enabled to the beginning of your developer message does not guarantee that the model will include markdown formatting in its response, it only increases the likelihood. We have found from internal testing that Formatting re-enabled is less effective by itself with the o1 model than with o3-mini.
To improve the performance of Formatting re-enabled you can further augment the beginning of the developer message which will often result in the desired output. Rather than just adding Formatting re-enabled to the beginning of your developer message, you can experiment with adding a more descriptive initial instruction like one of the examples below:
Formatting re-enabled - please enclose code blocks with appropriate markdown tags.Formatting re-enabled - code output should be wrapped in markdown.