Skip to main content
This feature is currently in public preview. This preview is provided without a service-level agreement, and is not recommended for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.
MAI-Transcribe is a next‑generation speech-to-text model built in‑house by the Microsoft AI team. It delivers fast and accurate transcription across 60 languages and real-world audio conditions. MAI-Transcribe-2 supports a wide range of workloads including video captioning, meetings, clinical notes, call center documentation, accessibility tools, content creation, and voice agents. The model provides speaker diarization, strong performance in noisy environments, word-level timestamps, automatic language identification, keyword biasing, code switching, and configurable transcription styles (clean transcripts without fillers, or verbatim). The following models are supported:
  • MAI-Transcribe-2
  • MAI-Transcribe-1.5
  • MAI-Transcribe-1: Deprecated on Aug 20, 2026.

Prerequisites

  • An Azure subscription. You can create one for free.
  • A Microsoft Foundry resource for Speech in the Azure portal.
  • The Speech resource key and region. After your Speech resource is deployed, select Go to resource to view and manage keys. For the current list of supported regions, see Speech service regions.
  • An audio file in one of these formats: WAV, MP3, or FLAC. For the maximum file size and audio duration, see the audio parameter in the Transcriptions - Transcribe REST reference. If you enable speaker diarization, see the note on recording length in the next section.

Use a MAI-Transcribe model

Use MAI‑Transcribe‑2 to generate transcripts from audio input. Configure the following features through the corresponding API parameters:

Language support

By default, the model operates in multilingual mode. The following languages are currently supported:

Use MAI-Transcribe with Voice Live

You can also use the MAI-Transcribe model for input audio transcription in the Voice Live API. Set the model field in the input_audio_transcription session configuration. For details, see How to customize Voice Live input and output.