Microsoft Introduces MAI-Transcribe-2 and MAI-Voice-2 Models for Speech AI

Microsoft has documented MAI-Transcribe-2, MAI-Voice-2, and MAI-Voice-2-Flash. The release expands its speech AI lineup, but native streaming transcription is not listed for MAI-Transcribe-2.

Microsoft Introduces MAI-Transcribe-2 and MAI-Voice-2 Models for Speech AI
Microsoft MAI-Transcribe-2 and MAI-Voice-2 Explained

Microsoft has introduced three documented MAI speech models: MAI-Transcribe-2 for speech-to-text, MAI-Voice-2 for high-fidelity text-to-speech, and MAI-Voice-2-Flash for lower-latency voice experiences. Together, they give developers and businesses distinct building blocks for transcribing audio, generating spoken responses, and designing more responsive voice-based interactions.

The most important distinction is that Microsoft’s official documentation does not list MAI-Transcribe-2-Streaming, MAI-Voice-2.1, or MAI-Voice-2.1-Flash as product names. It also does not advertise native real-time streaming transcription for MAI-Transcribe-2. The documented release is still significant, particularly for workflows that need multilingual transcription, speaker identification, timestamps, or more natural synthesized speech.

What Microsoft has documented

Microsoft’s official MAI-Transcribe-2 model catalog entry describes a multilingual speech-to-text model with diarization and word-level timestamps. The supplied documentation lists support for 60 languages, which can matter for companies processing customer calls, interviews, meetings, or other audio across multiple markets.

The roles of the three MAI speech models

The models address related but different parts of a speech workflow:

  • MAI-Transcribe-2 converts spoken audio to text. Its documented diarization capability is designed to distinguish speakers, while word-level timestamps can help teams locate a specific point in an audio recording.
  • MAI-Voice-2 converts text into high-fidelity speech. Microsoft positions it with broad language coverage and controls for prosody and emotion, which can affect how a generated voice delivers an answer.
  • MAI-Voice-2-Flash is a lower-latency version intended for real-time voice-agent scenarios, where reducing the delay before a spoken response can make a conversation feel more fluid.
Model Documented primary role Notable documented characteristics
MAI-Transcribe-2 Speech-to-text Multilingual support, diarization, word-level timestamps
MAI-Voice-2 Text-to-speech High-fidelity voice, broad language coverage, prosody and emotion control
MAI-Voice-2-Flash Lower-latency text-to-speech Optimized for real-time voice-agent interactions

Streaming transcription remains a key limitation

For businesses assessing call transcription or live agent-assist use cases, the missing documented streaming capability is important. MAI-Transcribe-2 may be suitable for processing recorded audio or workflows where transcription can occur after an audio segment is available. However, the supplied research indicates that native real-time streaming transcription is not yet publicly available in its documented feature set.

That does not diminish the value of speaker labels and word-level timestamps for post-call analysis, meeting notes, quality review, or searchable audio archives. It does mean teams should avoid designing a live transcription experience around an assumed MAI-Transcribe-2 streaming endpoint until Microsoft documents such support.

The catalog entry also notes pricing for a limited time, but the supplied research does not provide price amounts or terms. Businesses should therefore check the current Microsoft catalog and applicable Azure materials before estimating usage costs or committing to a production workflow.

Practical implications for speech automation

The release creates a clearer separation between transcription and voice response. A business can use transcription to turn spoken interactions into structured text, then use synthesized voice where spoken replies add value. Those are separate implementation decisions, with different latency and quality requirements.

Where the documented capabilities can fit

The documented features can support several practical patterns:

  • Turning recorded calls or interviews into searchable text with speaker attribution.
  • Sending timestamped transcripts to review, documentation, or follow-up workflows.
  • Generating spoken answers for automated phone or voice interfaces where a natural delivery matters.
  • Using the lower-latency Voice-2-Flash model when a voice agent needs to respond quickly between turns.

Diarization is especially useful when a transcript needs to show who said what. Word-level timestamps can make a long recording easier to audit because a reviewer can identify the precise point associated with a statement. Neither feature alone determines transcription accuracy for every use case, so teams should test representative audio, languages, accents, noise conditions, and conversation formats before relying on outputs in customer-facing or operational processes.

A sensible implementation approach

Businesses should begin with the workflow rather than the model name. For example, a recorded support call may need transcription, speaker separation, human review, and a transfer into a customer record. A voice agent may instead need a reliable application flow, a response-generation layer, and low-latency speech output.

That distinction also helps prevent a common planning error: treating low-latency speech synthesis as proof that the entire voice stack can operate in real time. MAI-Voice-2-Flash is documented for low-latency voice-agent use, while MAI-Transcribe-2 is not documented with native real-time streaming transcription. The overall experience will depend on how each component is connected and where delays occur.

Speech AI creates value only when transcription and voice outputs move reliably into the systems teams already use. Scalevise can design and implement AI automation workflows that route transcripts, preserve speaker context, trigger reviews, and connect approved outputs to operational tools. That can reduce manual handoffs while keeping people in control of customer-facing actions. Discuss an AI automation project with Scalevise.

Frequently Asked Questions

What is MAI-Transcribe-2?

MAI-Transcribe-2 is Microsoft’s documented multilingual speech-to-text model. The supplied research lists support for 60 languages, diarization, and word-level timestamps.

Does MAI-Transcribe-2 support native real-time streaming transcription?

No. The supplied research indicates that native real-time streaming transcription is not publicly available in the documented MAI-Transcribe-2 feature set.

What is the difference between MAI-Voice-2 and MAI-Voice-2-Flash?

MAI-Voice-2 is a high-fidelity text-to-speech model with prosody and emotion controls. MAI-Voice-2-Flash is a lower-latency variant optimized for real-time voice-agent interactions.

Is pricing available for the new MAI speech models?

The MAI-Transcribe-2 catalog entry notes pricing for a limited time, but the supplied research does not provide price amounts or terms. It also does not provide pricing details for the MAI voice models.


Conclusion

Microsoft’s documented MAI speech lineup provides distinct tools for transcription and voice generation, with useful capabilities such as diarization, word-level timestamps, expressive speech controls, and lower-latency voice output. The practical caveat is equally clear: MAI-Transcribe-2 is not currently documented as a native real-time streaming transcription model. Teams should match each model to a tested workflow and verify current pricing and product documentation before deployment.