# The Decoder — Microsoft AI releases new transcription and text-to-speech models for voice agents

- Company: The Decoder (the-decoder.com)
- Announced: 2026-10-02T09:20:39+00:00
- Category: new-model
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://the-decoder.com/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice-agents/
- Record: https://forck.live/items/15720-microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice
- Subject: Reporting
- Models affected: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, MAI-Voice-2.1-Flash
- Pricing: an hour of audio costs $0.54 at the introductory price; MAI-Voice-2.1-Flash costs $15 per million characters instead of $22

Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model that ranks first for accuracy on Artificial Analysis, transcribes 60 languages, and delivers first partial results in just over 100 milliseconds. It also released two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, which speak 23 languages in the same voice with a native accent in each. Both voice models can clone a voice from a few seconds of reference audio, and built-in safeguards are meant to prevent misuse.

## Evidence

Verbatim from https://the-decoder.com/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice-agents/:

> Microsoft says this lets voice agents respond while someone is still mid-sentence. Through the end of the year, an hour of audio costs $0.54 at the introductory price.

---

Record: https://forck.live/items/15720-microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
