Skip to main content
Use cloud evaluations to test generative AI applications at scale without managing local compute. This article sets up the shared SDK client and helps you choose a workflow for predeployment or production evaluation.

Prerequisites

  • A Foundry project.
  • An Azure OpenAI deployment with a GPT model that supports chat completion, such as gpt-5-mini.
  • The Foundry User role on the Foundry project.
The Foundry RBAC roles were recently renamed. Foundry User, Foundry Owner, Foundry Account Owner, and Foundry Project Manager were previously named Azure AI User, Azure AI Owner, Azure AI Account Owner, and Azure AI Project Manager. You might still see the previous names in some places while the rename rolls out. The role IDs and core permissions are unchanged by the rename.
Some evaluation features have regional restrictions. Review the supported regions before you begin.

Set up the SDK client

Install the SDK for your language. Set these shared environment variables:
  • AZURE_AI_PROJECT_ENDPOINT: Your Foundry project endpoint, for example, https://<account_name>.services.ai.azure.com/api/projects/<project_name>.
  • AZURE_AI_MODEL_DEPLOYMENT_NAME: The model deployment used by AI-assisted evaluators.
  • DATASET_NAME and DATASET_VERSION: Optional values for reusable datasets.
Authenticate with DefaultAzureCredential, create the project client, and get the OpenAI client used by the evaluation API.
Reference: AIProjectClient, DefaultAzureCredential
To use a model connected through admin connections as a target, judge model, or conversation simulator, see Use admin-connected models in cloud evaluations.

Understand the evaluation workflow

A cloud evaluation has three steps:
  1. Define the data shape and the evaluators that score it.
  2. Create the evaluation with the evaluation client.
  3. Start a run, poll until it completes, and retrieve the scored results.
Cloud evaluation results are stored in your Foundry project. You can retrieve them through the SDK, review them in the portal, or route them to Application Insights when it’s connected.

Choose evaluators

Evaluators bind to fields in your data through column mappings. Dataset workflows expose item fields, while target-generated workflows also expose the model or agent output through the sample schema. Review the built-in evaluators and custom evaluators before you configure testing criteria.

Choose your starting point

For adversarial safety testing, use AI red teaming. To create a standalone dataset, see Generate a synthetic evaluation dataset.