Microsoft introduced three new MAI speech models on Thursday, and the headline act is MAI-Transcribe-2-Streaming — a speech-to-text model designed not to wait for you to finish your sentence before it starts working.
Traditional transcription models operate in a listen-process-output loop. The streaming model instead produces an early draft the moment words arrive, keeps revising that text as more context flows in, and locks the final version once the speaker stops. Microsoft says in its internal tests text appears roughly twice as fast as the closest competitor in real-time dictation and captioning — a company-reported claim — and that the model can begin producing results just over 100 milliseconds after audio arrives.
The third-party numbers back the core claim. On Artificial Analysis' streaming speech recognition leaderboard, MAI-Transcribe-2-Streaming ranks first in accuracy with a final-transcript word error rate of 2.5%, returning final text about 0.13 seconds after the speaker stops; its first partial transcripts also score 2.5% WER at roughly 0.12 seconds of latency. That displaces SpaceXAI's Grok Voice Transcribe 2.0, the previous leader at 2.7% WER and 0.49 seconds.
The model supports 60 languages with automatic language detection, and is priced at $0.54 per hour of audio as an introductory rate through the end of 2026. The practical payoff is architectural: a customer-service agent can start diagnosing a caller's problem mid-sentence, and a voice assistant can begin planning tool calls while the user is still talking — shaving end-to-end latency out of the pipeline rather than the model.
Alongside it come two text-to-speech models: MAI-Voice-2.1 at $22 per million characters, aimed at higher-fidelity output, and MAI-Voice-2.1-Flash at $15 per million characters for latency- and cost-sensitive use cases. Both support 23 languages and 26 locales while keeping a consistent voice across languages. Microsoft AI CEO Mustafa Suleyman has also claimed the transcription model is 55% faster and 60% cheaper than ElevenLabs' offering — figures the company has not independently benchmarked.
The release extends Microsoft's push to build its own foundation models rather than relying solely on OpenAI. Earlier MAI speech models already power Copilot and Teams, and the new ones are exposed to developers through Microsoft Foundry — a signal that the internal models are graduating from internal plumbing to sellable product.
For the industry, the message is that voice is becoming a latency battleground. As agents take calls, join meetings, and act on spoken commands, whoever transcribes fastest and cheapest holds an unglamorous but compounding advantage — and for now, that is Microsoft, at about half a dollar per hour of audio.
Comments (0)
Log in to join the discussion
Log InNo comments yet