Skip to main content

Work with chat completion models

In this article, you send chat completion requests, build a multi-turn conversation, and manage the conversation’s token budget. Chat models are language models optimized for conversational interfaces. Unlike older text-in and text-out completion models, chat models accept a transcript of messages and return a model-generated message. This format supports multi-turn conversations and nonchat scenarios. Use the message format described in this article instead of prompting chat models like older completion models. Otherwise, the models might produce verbose or less useful responses.
For new apps, consider building on the Responses API instead of Chat Completions. To upgrade an existing app, see Azure OpenAI To Responses and Upgrade your Azure OpenAI app from Chat Completions to the Responses API.
Reasoning models, such as the GPT-5 series, behave differently on this API. They use max_completion_tokens instead of max_tokens, and they don’t support temperature, top_p, or the penalty parameters. On the gpt-5.6 and later models, a Chat Completions request that includes function tools fails unless you set reasoning_effort to none. Use the Responses API for tool calling with reasoning models. For details, see Azure OpenAI reasoning models.

Prerequisites

In the code samples, replace YOUR-RESOURCE-NAME with your Azure OpenAI resource name and YOUR-DEPLOYMENT-NAME with your model deployment name.