Skip to content
AI.info

The Pulse

Microsoft Ships Streaming Speech Models for Voice Agents

Microsoft launched MAI-Transcribe-2-Streaming, which returns live transcript updates in 60 languages, alongside two multilingual speech-generation models. The models are available through Microsoft Foundry and MAI Playground, with streaming

Microsoft Ships Streaming Speech Models for Voice Agents

AI.info Team ·

Microsoft’s new streaming transcription model is designed to give voice agents usable text before a caller finishes speaking. Announced on October 1, MAI-Transcribe-2-Streaming sends early transcript hypotheses just over 100 milliseconds after receiving audio, then updates them as more speech arrives and commits a stable version.

Partials let agents act mid-sentence

The model transcribes speech in 60 languages and automatically detects language as a conversation continues. Microsoft says those early transcript updates can let an agent begin reasoning or call a tool while the speaker is still talking, rather than waiting for a complete utterance.

For live dictation and subtitling, Microsoft says internal evaluations found that words appeared in its transcript twice as fast as with its closest competitor. The launch post does not name that competitor or detail the evaluation setup, so the comparison remains a company-reported result rather than a fully specified head-to-head test.

Microsoft says MAI-Transcribe-2-Streaming ranks first for accuracy on both final and partial transcripts in Artificial Analysis evaluations. Its earlier MAI-Transcribe-2 announcement covered a batch transcription model, not this streaming release: the company priced that model at $0.10 per audio hour, while the new streaming version carries an introductory rate of $0.54 per hour through the end of 2026.

Two voice models complete Microsoft’s audio set

Microsoft also introduced MAI-Voice-2.1 and MAI-Voice-2.1-Flash, both text-to-speech models. The standard version supports 23 languages across 26 locales, with Microsoft saying a single voice can retain its identity while speaking with a native accent in different languages; its listed price is $22 per million characters.

Flash targets high-volume, latency-sensitive uses such as spoken replies from voice agents. Microsoft says it can generate 45 seconds of audio with 150 milliseconds of end-to-end latency, offers 55% faster model inference, and costs about 60% less than comparable models; the listed price is $15 per million characters.

Both voice models can clone a voice across supported languages from a few seconds of reference audio. Microsoft says they include consent guardrails, but its announcement does not explain how those safeguards work or what checks customers must complete before using a cloned voice.

Where developers can try the models

Microsoft lists all three models in Microsoft Foundry and the MAI Playground, and says they are also available through Vercel and Azure Voice Live. OpenRouter offers the two voice models, while LiveKit support is listed as coming soon.

The launch pairs live transcription with speech generation, giving developers separate components for the listening and speaking sides of an agent. That can shorten delays between a user’s words and an agent’s response, but the posted timings and comparison claims are Microsoft’s; the release does not provide a complete end-to-end measurement for a working agent that also includes reasoning and tool calls.

Sources

Explore

More articles