Skip to main content
Items marked (preview) in this article are currently in public preview. This preview is provided without a service-level agreement, and we don’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
Model leaderboards (preview) in Foundry portal help you compare models in the Foundry model catalog using industry-standard model benchmarks. To get started, compare and select models using the model leaderboard in Foundry portal. You can review detailed benchmarking methodology for each leaderboard category:
  • Quality benchmarking of language models to understand how well models perform on core tasks including reasoning, knowledge, question answering, math, and coding.
  • Safety benchmarking of language models to understand how safe models are against harmful behavior generation.
  • Performance benchmarking of language models to understand how models perform in terms of latency and throughput.
  • Cost benchmarking of language models to understand the estimated cost of using models.
  • Scenario leaderboard benchmarking of language models to help you find the best model for your specific use case or scenario.
  • Quality benchmarking of embedding models to understand how well models perform on embedding-based tasks including search and retrieval.
When you find a suitable model, you can open its Detailed benchmarking results in the model catalog. From there, you can deploy the model, try it in the playground, or evaluate it on your own data. The leaderboards support benchmarking for text language models (including large language models (LLMs) and small language models (SLMs)) and embedding models. Model benchmarks assess LLMs and SLMs across quality, safety, cost, and throughput. Embedding models are evaluated using standard quality benchmarks. Leaderboards are updated as new models and benchmark datasets become available.

Model benchmarking scope

The model leaderboards feature a curated selection of text-based language models from the Foundry model catalog. Models are included based on the following criteria:
  • Foundry Models sold by Azure prioritized: Models sold by Azure are selected for relevance to common generative AI scenarios.
  • Core benchmark applicability: Models must support general-purpose language tasks such as reasoning, knowledge, question answering, mathematical reasoning, and coding. Specialized models (for example, protein folding or domain-specific QA) and other modalities aren’t supported.
This scoping ensures the leaderboards reflect current, high-quality models relevant to core AI scenarios.

Interpret leaderboard results

The leaderboards help you compare models across multiple dimensions so you can choose the right model for your use case. Here are some guidelines for interpreting the results:
  • Quality index: A higher quality index indicates stronger overall performance across reasoning, coding, math, and knowledge tasks. Compare the quality index across models to identify top performers for general-purpose language tasks.
  • Safety scores: Lower attack success rates indicate more robust models. Consider safety scores alongside quality scores, especially for customer-facing applications where harmful output is a significant concern.
  • Performance trade-offs: Use the latency and throughput metrics to understand the real-world responsiveness of a model. A model with high quality but high latency might not suit real-time applications.
  • Cost considerations: The estimated cost metric uses a three-to-one input-to-output token ratio. Adjust your expectations based on your actual workload’s input-to-output ratio.
  • Scenario leaderboards: If your use case maps to a specific scenario (for example, coding or math), start with the scenario leaderboard to find models optimized for that task rather than relying solely on the overall quality index.
Leaderboard benchmarks provide standardized comparisons across models using public datasets. To evaluate model performance on your specific data and use case, see Evaluate your generative AI apps.

Quality benchmarks of language models

Foundry assesses the quality of LLMs and SLMs using accuracy scores from standard benchmark datasets that measure reasoning, knowledge, question answering, math, and coding capabilities. Quality index values range from zero to one, where higher values indicate better performance. The datasets included in the quality index are: See more details in accuracy scores: Accuracy scores range from zero to one, where higher values are better.

Safety benchmarks of language models

Safety benchmarks are selected through a structured filtering and validation process designed to ensure both relevance and rigor. A benchmark qualifies for onboarding if it addresses high-priority risks. The safety leaderboards include benchmarks that are reliable enough to provide meaningful signals on topics of interest as they relate to safety. The leaderboards use HarmBench to proxy model safety, and organize scenario leaderboards as follows:

Harmful behavior detection

The HarmBench benchmark measures harmful behaviors using prompts designed to elicit unsafe responses. It covers seven semantic categories:
  • Cybercrime and unauthorized intrusion
  • Chemical and biological weapons or drugs
  • Copyright violations
  • Misinformation and disinformation
  • Harassment and bullying
  • Illegal activities
  • General harm
These categories are grouped into three functional areas:
  • Standard harmful behaviors
  • Contextually harmful behaviors
  • Copyright violations
Each functional category is featured in a separate scenario leaderboard. The evaluation uses direct prompts from HarmBench (no attacks) and HarmBench evaluators to calculate Attack Success Rate (ASR). Lower ASR values mean safer models. No attack strategies are used for evaluation, and model benchmarking is performed with Foundry Guardrails (previously content filters) turned off.

Toxic content detection

Toxigen is a large-scale dataset for detecting adversarial and implicit hate speech. It includes implicitly toxic and benign sentences referencing 13 minority groups. Foundry uses annotated Toxigen samples and calculates F1 scores to measure classification performance. Higher scores indicate better toxic content detection. Benchmarking is performed with Foundry Guardrails (previously content filters) turned off.

Sensitive domain knowledge

The Weapons of Mass Destruction Proxy (WMDP) benchmark measures model knowledge in sensitive domains including biosecurity, cybersecurity, and chemical security. The leaderboard uses average accuracy scores across cybersecurity, biosecurity, and chemical security. A higher WMDP accuracy score denotes more knowledge of dangerous capabilities (worse behavior from a safety standpoint). Model benchmarking is performed with the default Foundry Guardrails (previously content filters) on. These guardrails detect and block content harm in violence, self-harm, sexual, hate, and unfairness, but don’t target categories in cybersecurity, biosecurity, and chemical security.

Limitations of safety benchmarks

Safety is a complex topic with several dimensions. No single open-source benchmark can test or represent the full safety of a system across all scenarios. Additionally, many benchmarks suffer from saturation or misalignment between benchmark design and risk definition. Some benchmarks also lack clear documentation on how targets risks are conceptualized and operationalized, making it difficult to assess whether results accurately capture the nuances of real-world risks. These limitations can lead to either overestimating or underestimating model performance in real-world safety scenarios.

Performance benchmarks of language models

Performance metrics are aggregated over 14 days using 24 trials per day, with two requests per trial sent at one-hour intervals. Unless otherwise noted, the following default parameters apply to both serverless API deployments and Azure OpenAI: The performance of LLMs and SLMs is assessed across the following metrics: Foundry summarizes performance using: For performance metrics like latency or throughput, the time to first token and the generated tokens per second give a better overall sense of the typical performance and behavior of the model. Performance numbers are refreshed periodically to reflect the latest deployment configurations.

Cost benchmarks of language models

Cost benchmarks measure the actual cost to execute each model on the quality benchmark datasets, rather than an estimated cost based on token pricing. The benchmark cost is computed using:
  • Actual number of input, reasoning, and output tokens consumed during benchmark execution.
  • Model-specific reasoning effort configuration used for evaluation (typically high or xhigh).
  • Dataset characteristics and complexity, which affect token usage and runtime.
Unlike estimates based on a fixed token ratio, this approach reflects the true end-to-end cost of running the benchmark workloads.

How to interpret cost results

  • Cost is reported in USD per benchmark run across the standard quality datasets.
  • Values represent real execution cost and enable direct comparison between models.
  • Lower values indicate more cost-efficient performance on the benchmark suite.

Scenario leaderboard benchmarking

Scenario leaderboards group benchmark datasets by common real-world evaluation goals. You can quickly identify a model’s strengths and weaknesses by use case. Each scenario aggregates one or more public benchmark datasets. Use the following table to find your use case in the Scenario column, and then review the associated benchmark datasets and what the results indicate. The following table summarizes the available scenario leaderboards and their associated datasets and descriptions:

Quality benchmarks of embedding models

The quality index of embedding models is defined as the averaged accuracy scores of a comprehensive set of serverless API benchmark datasets targeting Information Retrieval, Document Clustering, and Summarization tasks.

Calculation of scores

Individual scores

Benchmark results originate from public datasets that are commonly used for language model evaluation. In most cases, the data is hosted in GitHub repositories maintained by the creators or curators of the data. Foundry evaluation pipelines download data from their original sources, extract prompts from each example row, generate model responses, and then compute relevant accuracy metrics. Prompt construction follows best practices for each dataset, as specified by the paper introducing the dataset and industry standards. In most cases, each prompt contains several shots, that is, several examples of complete questions and answers to prime the model for the task. The number of shots varies by dataset and follows the methodology specified in each dataset’s original publication. The evaluation pipelines create shots by sampling questions and answers from a portion of the data held out from evaluation.

Benchmark limitations

All benchmarks have inherent limitations that you should consider when interpreting results:
  • Quality benchmarks: Benchmark datasets can become saturated over time as models are trained or tuned on similar data. Evaluation results might also vary depending on prompt construction and the number of few-shot examples used.
  • Performance benchmarks: Metrics are collected using synthetic workloads with a fixed input-to-output token ratio and single-region deployments. Real-world performance might differ based on workload patterns, concurrency, region, and deployment configuration.
  • Cost benchmarks: Cost estimates are based on a three-to-one input-to-output token ratio and current pricing at the time of measurement. Actual costs depend on your workload and are subject to pricing changes.