Items marked preview in this article are currently in preview. This preview is provided without a service-level agreement, and Microsoft doesn’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
- Simple QnA (single-turn) produces question-and-answer pairs for turn-level evaluation.
- Simulation seed (multi-turn) produces scenario descriptions that feed the Simulate conversations flow for multi-turn evaluation.
- Agent definition—seed generation from a deployed agent’s instructions or prompt.
- Prompt—pass an inline text prompt that describes the domain or steers difficulty.
- Reference file—upload a document (for example, a policy, spec, or knowledge-base export) and generate questions grounded in its content.
When to use synthetic generation
Use synthetic generation when:- You’re prelaunch and have no production traces yet.
- Your agent has low traffic and a trace window doesn’t yield enough distinct samples.
- You need a stable regression baseline that doesn’t drift with changing production behavior.
- You want to expand coverage of edge cases the agent hasn’t encountered in production.
- You’re iterating on an agent’s instructions and want a quick smoke-test dataset.
Choose a source type
You can combine sources in a single job. A common pattern is to pair a reference file (for grounding) with a prompt (for steering tone or difficulty).
Prerequisites
- Python SDK version
2.5.0or later:pip install "azure-ai-projects>=2.5.0" azure-identity. - A Microsoft Foundry project endpoint URL in the format
https://<your-resource>.services.ai.azure.com/api/projects/<your-project>. - Foundry User role or higher on the project.
- For all evaluation role requirements, see Set up permissions for evaluation workflows.
- An Azure OpenAI model deployment that supports the Responses API. Both the
simple_qnaandsimulation_seedrecipes use this model to synthesize output rows. For the supported-model list, see Azure OpenAI Responses API model support. - A supported region. For the list, see Supported regions for data generation.
Generate a dataset from the portal
- In the portal, open the Data Generation tab. Select Create dataset, and then select Generate synthetic.
- In Generate synthetic data, set Dataset usage to Evaluation.
- Set Task type. Select Simple QnA (single-turn) for question-and-answer pairs, or Simulation seed (multi-turn) for scenario descriptions used in conversation simulation.
- Select a Generator model.
- Provide one or more source inputs: Agent, Prompt, or Reference file.
- Set Maximum number of samples and Output file name.
- Select Generate.
- Track the dataset generation job status in the Data Generation tab.
- When the job finishes, preview the generated rows on the Data tab.

Generate a dataset from an agent definition (SDK)
This flow seeds generation from a deployed agent’s instructions. The service fetches the agent’s prompt and uses your configured model to synthesize question-and-answer pairs from it. First, create anAIProjectClient by using your project endpoint and DefaultAzureCredential. You can find all data generation operations under project_client.beta.datasets.
- Python
- JavaScript/TypeScript
SimpleQnA job whose source is an agent reference. If you already have a deployed agent, skip the create_version call and pass its existing name and version to AgentDataGenerationJobSource.
query and ground_truth fields. Preview it on the Data tab in the portal to spot-check the generated rows before evaluating.
Generate a dataset from a prompt (SDK)
If you don’t have a deployed agent yet, or if you want to generate data from a self-contained snippet of source material, pass the text as aPromptDataGenerationJobSource. This approach is useful for policy documents, FAQ content, or short specs.
- Python
- JavaScript/TypeScript
Generate a dataset from reference files (SDK)
For longer source material, upload a document as an Azure OpenAI file and reference it by ID. This option works best when the agent’s domain knowledge lives in a spec, knowledge-base export, or policy document, because the generated questions stay grounded in that content. The file must be in theprocessed state before the data generation service can use it, and it needs to contain at least 1 KB of content.
Supported reference file extensions are:
.txt, .md, .csv, .json, .xml, .html, .pdf, .png, .jpg, .jpeg, .gif, .tiff, .tif, .svg.
Generate a simulation seed dataset (SDK)
Simulation seed jobs produce a dataset of scenario descriptions that feed the Simulate conversations flow. Generated rows can containid, category, test_case_description, and desired_num_turns. Only test_case_description is required.
The job shape is identical to Simple Q&A. The only differences are the options class (SimulationSeedDataGenerationJobOptions) and the wire type value (simulation_seed). The following example uses an agent definition as the source. To use a prompt or reference file instead, swap the source class as shown in Generate a dataset from a prompt (SDK) or Generate a dataset from reference files (SDK), and substitute SimulationSeedDataGenerationJobOptions for SimpleQnADataGenerationJobOptions.
This example assumes a deployed agent named retail-agent. If you don’t have one yet, create it first with the create_version pattern shown in Generate a dataset from an agent definition (SDK).
result by using the output-handling pattern shown in Generate a dataset from an agent definition (SDK).
Generated dataset schema
The simulation seed schema supports the following fields. Onlytest_case_description is required; the other fields are optional.
Example row (
test_case_description shortened for readability):
test_case_description values tend to produce lower-quality simulated conversations.
Run an evaluation against the generated dataset
The evaluation path depends on the task type:- Simple Q&A datasets use the standard
queryandground_truthschema and work directly with the evaluation APIs. For the full flow, see Evaluate models and agents in the cloud. For complete runnable end-to-end examples, see sample_synthetic_data_agent_evaluation.py and sample_synthetic_data_model_evaluation.py on GitHub. - Simulation seed datasets feed the Simulate conversations flow. Pass the generated dataset ID as the simulation run’s source. The simulator uses each row’s
test_case_descriptionto play the user’s role and, when provided, usesdesired_num_turnsas guidance. Conversation-level evaluators score the resulting conversation rather than the seed row.
Manage data generation jobs
Useproject_client.beta.datasets job-management APIs to list, inspect, cancel, and delete synthetic generation jobs.
- Python
- JavaScript/TypeScript
Limitations
- If your Foundry project is connected to your own storage account, public network access must be enabled on that storage account for successful dataset creation.
Best practices
- Mirror your production system prompt. When you generate from an agent definition or a prompt, use instructions that match what your production agent actually runs. Drift here weakens the evaluation signal.
- Combine a reference file with a prompt for grounded coverage. The file anchors generated questions in real domain content; the prompt steers tone, difficulty, or topic emphasis.
- Generate a small batch first. Start at the minimum
max_samplesof 15, review the rows manually on the Data tab, then scale up once the output quality looks right. - Regenerate when the agent’s instructions change. A dataset generated from one version of an agent’s prompt becomes stale when the prompt changes significantly. Rerun the job and version the new output.
- Combine synthetic and trace-based generation for the strongest coverage. Synthetic data fills gaps before launch and for edge cases; production traces reflect how your agent actually behaves. Use both sources together rather than treating them as alternatives. See Convert agent traces into evaluation datasets.
- Write scenario-focused
test_case_descriptionvalues for simulation seeds. The simulator plays the user side of the conversation based on this text. Descriptions that spell out the user’s goal, constraints, and any edge cases you want to cover produce higher-quality simulated conversations.