Skip to main content
Items marked preview in this article are currently in preview. This preview is provided without a service-level agreement, and Microsoft doesn’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
This article covers synthetic data generation. For all dataset preparation options and the standard field names, see Evaluation datasets in Microsoft Foundry and Evaluation dataset schema. When your agent doesn’t have production traffic yet, you can still build a meaningful evaluation dataset. The Microsoft Foundry data generation service synthesizes evaluation data from material you already have: an agent’s instructions, an inline prompt, or a reference document you upload. Two task types are available:
  • Simple QnA (single-turn) produces question-and-answer pairs for turn-level evaluation.
  • Simulation seed (multi-turn) produces scenario descriptions that feed the Simulate conversations flow for multi-turn evaluation.
You can evaluate simple Q&A datasets directly. Simulation seed datasets first drive a simulator that plays the user’s role against the target agent. Conversation-level evaluators then score the generated conversations. Three input source types are available, and you can combine them in a single job for richer coverage:
  • Agent definition—seed generation from a deployed agent’s instructions or prompt.
  • Prompt—pass an inline text prompt that describes the domain or steers difficulty.
  • Reference file—upload a document (for example, a policy, spec, or knowledge-base export) and generate questions grounded in its content.
Synthetic generation and trace-based generation are complementary: synthetic datasets cover edge cases and prelaunch scenarios, while trace-based datasets reflect real production behavior. Using both gives the strongest evaluation signal. See Convert agent traces into evaluation datasets.

When to use synthetic generation

Use synthetic generation when:
  • You’re prelaunch and have no production traces yet.
  • Your agent has low traffic and a trace window doesn’t yield enough distinct samples.
  • You need a stable regression baseline that doesn’t drift with changing production behavior.
  • You want to expand coverage of edge cases the agent hasn’t encountered in production.
  • You’re iterating on an agent’s instructions and want a quick smoke-test dataset.

Choose a source type

You can combine sources in a single job. A common pattern is to pair a reference file (for grounding) with a prompt (for steering tone or difficulty).

Prerequisites

  • Python SDK version 2.5.0 or later: pip install "azure-ai-projects>=2.5.0" azure-identity.
  • A Microsoft Foundry project endpoint URL in the format https://<your-resource>.services.ai.azure.com/api/projects/<your-project>.
  • Foundry User role or higher on the project.
  • For all evaluation role requirements, see Set up permissions for evaluation workflows.
  • An Azure OpenAI model deployment that supports the Responses API. Both the simple_qna and simulation_seed recipes use this model to synthesize output rows. For the supported-model list, see Azure OpenAI Responses API model support.
  • A supported region. For the list, see Supported regions for data generation.

Generate a dataset from the portal

  1. In the portal, open the Data Generation tab. Select Create dataset, and then select Generate synthetic.
  2. In Generate synthetic data, set Dataset usage to Evaluation.
  3. Set Task type. Select Simple QnA (single-turn) for question-and-answer pairs, or Simulation seed (multi-turn) for scenario descriptions used in conversation simulation.
  4. Select a Generator model.
  5. Provide one or more source inputs: Agent, Prompt, or Reference file.
  6. Set Maximum number of samples and Output file name.
  7. Select Generate.
  8. Track the dataset generation job status in the Data Generation tab.
  9. When the job finishes, preview the generated rows on the Data tab.
Screenshot of the Generate synthetic data dialog showing Dataset usage set to Evaluation, Task type set to Simple Q&A, Generator model, source inputs, Maximum number of samples, and Output file name.

Generate a dataset from an agent definition (SDK)

This flow seeds generation from a deployed agent’s instructions. The service fetches the agent’s prompt and uses your configured model to synthesize question-and-answer pairs from it. First, create an AIProjectClient by using your project endpoint and DefaultAzureCredential. You can find all data generation operations under project_client.beta.datasets.
Then submit a SimpleQnA job whose source is an agent reference. If you already have a deployed agent, skip the create_version call and pass its existing name and version to AgentDataGenerationJobSource.
The job produces a versioned dataset with single-turn query and ground_truth fields. Preview it on the Data tab in the portal to spot-check the generated rows before evaluating.

Generate a dataset from a prompt (SDK)

If you don’t have a deployed agent yet, or if you want to generate data from a self-contained snippet of source material, pass the text as a PromptDataGenerationJobSource. This approach is useful for policy documents, FAQ content, or short specs.
Resolve the dataset by using the same pattern shown in the previous section.

Generate a dataset from reference files (SDK)

For longer source material, upload a document as an Azure OpenAI file and reference it by ID. This option works best when the agent’s domain knowledge lives in a spec, knowledge-base export, or policy document, because the generated questions stay grounded in that content. The file must be in the processed state before the data generation service can use it, and it needs to contain at least 1 KB of content. Supported reference file extensions are: .txt, .md, .csv, .json, .xml, .html, .pdf, .png, .jpg, .jpeg, .gif, .tiff, .tif, .svg.

Generate a simulation seed dataset (SDK)

Simulation seed jobs produce a dataset of scenario descriptions that feed the Simulate conversations flow. Generated rows can contain id, category, test_case_description, and desired_num_turns. Only test_case_description is required. The job shape is identical to Simple Q&A. The only differences are the options class (SimulationSeedDataGenerationJobOptions) and the wire type value (simulation_seed). The following example uses an agent definition as the source. To use a prompt or reference file instead, swap the source class as shown in Generate a dataset from a prompt (SDK) or Generate a dataset from reference files (SDK), and substitute SimulationSeedDataGenerationJobOptions for SimpleQnADataGenerationJobOptions. This example assumes a deployed agent named retail-agent. If you don’t have one yet, create it first with the create_version pattern shown in Generate a dataset from an agent definition (SDK).
Resolve the generated dataset from result by using the output-handling pattern shown in Generate a dataset from an agent definition (SDK).

Generated dataset schema

The simulation seed schema supports the following fields. Only test_case_description is required; the other fields are optional. Example row (test_case_description shortened for readability):
Preview the generated rows on the Data tab before running conversation simulation. Rows with unclear test_case_description values tend to produce lower-quality simulated conversations.

Run an evaluation against the generated dataset

The evaluation path depends on the task type:

Manage data generation jobs

Use project_client.beta.datasets job-management APIs to list, inspect, cancel, and delete synthetic generation jobs.
For more context, see Manage data generation jobs.

Limitations

  • If your Foundry project is connected to your own storage account, public network access must be enabled on that storage account for successful dataset creation.

Best practices

  • Mirror your production system prompt. When you generate from an agent definition or a prompt, use instructions that match what your production agent actually runs. Drift here weakens the evaluation signal.
  • Combine a reference file with a prompt for grounded coverage. The file anchors generated questions in real domain content; the prompt steers tone, difficulty, or topic emphasis.
  • Generate a small batch first. Start at the minimum max_samples of 15, review the rows manually on the Data tab, then scale up once the output quality looks right.
  • Regenerate when the agent’s instructions change. A dataset generated from one version of an agent’s prompt becomes stale when the prompt changes significantly. Rerun the job and version the new output.
  • Combine synthetic and trace-based generation for the strongest coverage. Synthetic data fills gaps before launch and for edge cases; production traces reflect how your agent actually behaves. Use both sources together rather than treating them as alternatives. See Convert agent traces into evaluation datasets.
  • Write scenario-focused test_case_description values for simulation seeds. The simulator plays the user side of the conversation based on this text. Descriptions that spell out the user’s goal, constraints, and any edge cases you want to cover produce higher-quality simulated conversations.