Work with chat completion models
In this article, you send chat completion requests, build a multi-turn conversation, and manage the conversation’s token budget. Chat models are language models optimized for conversational interfaces. Unlike older text-in and text-out completion models, chat models accept a transcript of messages and return a model-generated message. This format supports multi-turn conversations and nonchat scenarios. Use the message format described in this article instead of prompting chat models like older completion models. Otherwise, the models might produce verbose or less useful responses.Reasoning models, such as the GPT-5 series, behave differently on this API. They use
max_completion_tokens instead of max_tokens, and they don’t support temperature, top_p, or the penalty parameters. On the gpt-5.6 and later models, a Chat Completions request that includes function tools fails unless you set reasoning_effort to none. Use the Responses API for tool calling with reasoning models. For details, see Azure OpenAI reasoning models.Prerequisites
- An Azure OpenAI resource with a chat completion model deployment. To create a resource and deploy a model, see Create a resource and deploy a model with Azure OpenAI.
YOUR-RESOURCE-NAME with your Azure OpenAI resource name and YOUR-DEPLOYMENT-NAME with your model deployment name.