DeepSeek-V4-Pro, by using the OpenAI Chat Completions API from a Foundry project.
Prerequisites
- An Azure subscription.
- A Foundry project. This kind of project is managed under a Foundry resource. If you don’t have a Foundry project, see Create a project for Microsoft Foundry.
- Your Foundry project’s endpoint URL, which is of the form
https://YOUR-RESOURCE-NAME.services.ai.azure.com/api/projects/YOUR_PROJECT_NAME. - A deployed reasoning model. This article uses
DeepSeek-V4-Pro; replace the model name with your deployment name when necessary. - The SDK or command-line tools for the language you select. For C#, use the .NET 10 SDK. The C# examples are tested with the
OpenAI2.13.0 andAzure.Identity1.21.0 packages. - Permission to access the project and model deployment. For keyless authentication, sign in with an identity that has the required Foundry project access.
Use the AI model starter kit
The examples in the AI model starter kit use standard OpenAI clients with a Foundry project endpoint and the/openai/v1 path. The starter kit includes a complete example for a DeepSeek reasoning model. Its current samples use the Responses API. The examples in this article use Chat Completions for deployments that expose that API.
Set up the client
Use Microsoft Entra ID to authenticate the OpenAI client. The token scope for the Foundry project endpoint ishttps://ai.azure.com/.default. The endpoint must be the project endpoint, not the resource endpoint, and the client base URL must append /openai/v1.
Create a basic chat completion
Send a user message to the deployed reasoning model. The response contains the final answer inmessage.content. Depending on the model, the response can also contain reasoning content in message.reasoning_content.
Read reasoning content
Some non-OpenAI reasoning models return areasoning_content field alongside the final content. Reasoning content can be lengthy and counts toward token usage. Don’t add it to the message history for a multi-turn conversation unless the model documentation specifically requires it. Store or display it only when your application needs it, and treat it as model output rather than a verified explanation.
For models that don’t return reasoning_content, use the final content value. The field is model-dependent and isn’t available for every reasoning model.
Stream a completion
Setstream to true to receive server-sent events as the model generates output. Reasoning content and final answer content can arrive in different delta fields. The following examples print final answer content as it arrives.
Choose parameters for reasoning models
Reasoning models often don’t support parameters that are common for other chat completion models, includingtemperature, top_p, presence_penalty, and frequency_penalty. Check the model details in the Foundry model catalog before adding optional parameters. Set a sufficient max_completion_tokens value because reasoning tokens and final answer tokens both count toward the limit.
Use short, direct prompts. Avoid asking the model to reveal a chain of thought. For multi-turn conversations, append the final answer instead of the reasoning content when the model returns both.
Handle content safety responses
Foundry applies content filtering to supported deployments. A request or response can be blocked when it violates a configured content safety policy. Handle thecontent_filter finish reason and HTTP 400 errors in your application, show a useful message to the user, and don’t retry the same blocked prompt unchanged.
About reasoning models
Reasoning models reach higher levels of performance in domains like math, coding, science, strategy, and logistics. These models explicitly use a chain of thought to explore all possible paths before generating an answer. They verify their answers as they produce them, which helps them arrive at more accurate conclusions. As a result, reasoning models might require fewer context prompts to produce effective results. Reasoning models produce two types of content as outputs:- Reasoning completions
- Output completions
DeepSeek-V4-Pro, might respond with the reasoning content. Others, like o1, output only the completions.
Prompt reasoning models
When building prompts for reasoning models, take the following into consideration:- Use simple instructions and avoid using chain-of-thought techniques.
- Built-in reasoning capabilities make simple zero-shot prompts as effective as more complex methods.
- When providing additional context or documents, like in RAG scenarios, including only the most relevant information might help prevent the model from over-complicating its response.
- Reasoning models may support the use of system messages. However, they might not follow them as strictly as other non-reasoning models.
- When creating multi-turn applications, consider appending only the final answer from the model, without its reasoning content. Notice that reasoning models can take longer times to generate responses. They use long reasoning chains of thought that enable deeper and more structured problem-solving. They also perform self-verification to cross-check their answers and correct their mistakes, thereby showcasing emergent self-reflective behaviors.