Prerequisites
- A Foundry project.
- An Azure OpenAI deployment with a GPT model that supports chat completion, such as
gpt-5-mini. - The Foundry User role on the Foundry project.
The Foundry RBAC roles were recently renamed. Foundry User, Foundry Owner, Foundry Account Owner, and Foundry Project Manager were previously named Azure AI User, Azure AI Owner, Azure AI Account Owner, and Azure AI Project Manager. You might still see the previous names in some places while the rename rolls out. The role IDs and core permissions are unchanged by the rename.
- For all evaluation role requirements, see Set up permissions for evaluation workflows.
- Optionally, your own storage account for evaluation data.
Set up the SDK client
Install the SDK for your language. Set these shared environment variables:AZURE_AI_PROJECT_ENDPOINT: Your Foundry project endpoint, for example,https://<account_name>.services.ai.azure.com/api/projects/<project_name>.AZURE_AI_MODEL_DEPLOYMENT_NAME: The model deployment used by AI-assisted evaluators.DATASET_NAMEandDATASET_VERSION: Optional values for reusable datasets.
DefaultAzureCredential, create the project client, and get the OpenAI client used by the evaluation API.
- Python
- C#
- JavaScript/TypeScript
Understand the evaluation workflow
A cloud evaluation has three steps:- Define the data shape and the evaluators that score it.
- Create the evaluation with the evaluation client.
- Start a run, poll until it completes, and retrieve the scored results.
Choose evaluators
Evaluators bind to fields in your data through column mappings. Dataset workflows expose item fields, while target-generated workflows also expose the model or agent output through the sample schema. Review the built-in evaluators and custom evaluators before you configure testing criteria.Choose your starting point
For adversarial safety testing, use AI red teaming. To create a standalone dataset, see Generate a synthetic evaluation dataset.