Skip to main content
This example uses the OpenAI-compatible Chat Completions API through a Foundry Models endpoint. Function calling through the Responses API uses a different request and response format. For more information, see Use the Azure OpenAI Responses API.
If you include one or more functions in your request, the model decides whether to call any of the functions based on the context of the prompt. When the model decides to call a function, it responds with a JSON object that includes the arguments for the function. The model generates function calls and structured arguments based on the functions you specify. Your application executes the functions and returns the results to the model, so you remain in control of the actions taken. At a high level, you can break down working with functions into three steps:
  1. Call the chat completions API with your functions and the user’s input
  2. Use the model’s response to call your API or function
  3. Call the chat completions API again, including the response from your function to get a final response

Prerequisites

  • A Foundry Model deployment that supports function calling through the Chat Completions API.
  • A Foundry resource endpoint in the format https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/.
  • For Microsoft Entra ID authentication:
    • A custom subdomain configured. For more information, see Custom subdomain names.
    • Packages: pip install openai azure-identity.

Single tool/function calling example

This example demonstrates a toy function call that checks the time in three hardcoded locations with a single tool/function defined. Print statements make the code execution easier to follow:
Output:
If we’re using a model deployment that supports parallel function calls we could convert this into a parallel function calling example by changing the messages array to ask for the time in multiple locations instead of one. To accomplish this swap the comments in these two lines:
To look like this, and run the code again:
This generates the following output: Output:
Parallel function calls allow you to perform multiple function calls together, allowing for parallel execution and retrieval of results. This reduces the number of calls to the API that need to be made and can improve overall performance. For example in our simple time app we retrieved multiple times at the same time. This resulted in a chat completion message with three function calls in the tool_calls array, each with a unique id. If you wanted to respond to these function calls, you would add three new messages to the conversation, each containing the result of one function call, with a tool_call_id referencing the id from tool_calls. To force the model to call a specific function, set tool_choice to a named tool object, for example: tool_choice={"type": "function", "function": {"name": "get_current_time"}}. To force a user-facing message, set tool_choice="none".
The default behavior (tool_choice: "auto") is for the model to decide on its own if it should call a function and if so which function to call.

Parallel function calling with multiple functions

Now we demonstrate another toy function calling example this time with two different tools/functions defined.
Output
The JSON response might not always be valid so you need to add additional logic to your code to be able to handle errors. For some use cases you may find you need to use fine-tuning to improve function calling performance.

Prompt engineering with functions

When you define a function as part of your request, the details are injected into the system message using specific syntax that the model has been trained on. This means that functions consume tokens in your prompt and that you can apply prompt engineering techniques to optimize the performance of your function calls. The model uses the full context of the prompt to determine if a function should be called including function definition, the system message, and the user messages.

Improving quality and reliability

If the model isn’t calling your function when or how you expect, there are a few things you can try to improve the quality.
Provide more details in your function definition
It’s important that you provide a meaningful description of the function and provide descriptions for any parameter that might not be obvious to the model. For example, in the description for the location parameter, you could include extra details and examples on the format of the location.
Provide more context in the system message
The system message can also be used to provide more context to the model. For example, if you have a function called search_hotels you could include a system message like the following to instruct the model to call the function when a user asks for help with finding a hotel.
Instruct the model to ask clarifying questions
In some cases, you want to instruct the model to ask clarifying questions to prevent making assumptions about what values to use with functions. For example, with search_hotels you would want the model to ask for clarification if the user request didn’t include details on location. To instruct the model to ask a clarifying question, you could include content like the next example in your system message.

Reducing errors

Another area where prompt engineering can be valuable is in reducing errors in function calls. The models are trained to generate function calls matching the schema that you define, but the models produce a function call that doesn’t match the schema you defined or try to call a function that you didn’t include. If you find the model is generating function calls that weren’t provided, try including a sentence in the system message that says "Only use the functions you have been provided with.".

Using function calling responsibly

Like any AI system, using function calling to integrate language models with other tools and systems presents potential risks. It’s important to understand the risks that function calling could present and take measures to ensure you use the capabilities responsibly. Here are a few tips to help you use functions safely and securely:
  • Validate Function Calls: Always verify the function calls generated by the model. This includes checking the parameters, the function being called, and ensuring that the call aligns with the intended action.
  • Use Trusted Data and Tools: Only use data from trusted and verified sources. Untrusted data in a function’s output could be used to instruct the model to write function calls in a way other than you intended.
  • Follow the Principle of Least Privilege: Grant only the minimum access necessary for the function to perform its job. This reduces the potential impact if a function is misused or exploited. For example, if you’re using function calls to query a database, you should only give your application read-only access to the database. You also shouldn’t depend solely on excluding capabilities in the function definition as a security control.
  • Consider Real-World Impact: Be aware of the real-world impact of function calls that you plan to execute, especially those that trigger actions such as executing code, updating databases, or sending notifications.
  • Implement User Confirmation Steps: Particularly for functions that take actions, we recommend including a step where the user confirms the action before execution.
To learn more about responsible AI practices for Foundry Models, see the Overview of responsible AI practices.
In the Chat Completions API, the functions and function_call parameters are deprecated. Use the tools parameter instead of functions, and use the tool_choice parameter instead of function_call.

Function calling support

Foundry Models support function calling through the Chat Completions API or the Responses API. Support varies by model, API, deployment type, and model version. To find a model that supports function calling, review the following catalogs: In the model details, look for the following capabilities:
  • Chat Completions API or Responses API, depending on the API you plan to use.
  • Functions or tools for function calling support.
  • Parallel tool calling if your application needs the model to request multiple function calls in one response.
This article shows how to implement function calling with the Chat Completions API. For the Responses API request and response format, see Use the Azure OpenAI Responses API.
The tool_choice parameter is now supported with o3-mini and o1. For more information, see the reasoning models guide.
Tool and function descriptions are currently limited to 1,024 characters. We’ll update this article if this limit changes.

Next steps