Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Microsoft's Streaming Speech Model Starts Writing While You're Still Talking — and Tops the Accuracy Board at 2.5% WER, Beating Grok

Microsoft's Streaming Speech Model Starts Writing While You're Still Talking — and Tops the Accuracy Board at 2.5% WER, Beating Grok

Microsoft launched three MAI speech models led by MAI-Transcribe-2-Streaming, which ranks first on Artificial Analysis' streaming speech recognition test with a 2.5% word error rate, ahead of Grok Voice Transcribe 2.0 at 2.7%. Microsoft says words appear roughly 100 milliseconds after audio arrives (company-reported). Intro pricing is $0.54 per hour through year-end, across 60 languages.

Microsoft introduced three new MAI speech models on Thursday, and the headline act is MAI-Transcribe-2-Streaming — a speech-to-text model designed not to wait for you to finish your sentence before it starts working.

Traditional transcription models operate in a listen-process-output loop. The streaming model instead produces an early draft the moment words arrive, keeps revising that text as more context flows in, and locks the final version once the speaker stops. Microsoft says in its internal tests text appears roughly twice as fast as the closest competitor in real-time dictation and captioning — a company-reported claim — and that the model can begin producing results just over 100 milliseconds after audio arrives.

The third-party numbers back the core claim. On Artificial Analysis' streaming speech recognition leaderboard, MAI-Transcribe-2-Streaming ranks first in accuracy with a final-transcript word error rate of 2.5%, returning final text about 0.13 seconds after the speaker stops; its first partial transcripts also score 2.5% WER at roughly 0.12 seconds of latency. That displaces SpaceXAI's Grok Voice Transcribe 2.0, the previous leader at 2.7% WER and 0.49 seconds.

The model supports 60 languages with automatic language detection, and is priced at $0.54 per hour of audio as an introductory rate through the end of 2026. The practical payoff is architectural: a customer-service agent can start diagnosing a caller's problem mid-sentence, and a voice assistant can begin planning tool calls while the user is still talking — shaving end-to-end latency out of the pipeline rather than the model.

Alongside it come two text-to-speech models: MAI-Voice-2.1 at $22 per million characters, aimed at higher-fidelity output, and MAI-Voice-2.1-Flash at $15 per million characters for latency- and cost-sensitive use cases. Both support 23 languages and 26 locales while keeping a consistent voice across languages. Microsoft AI CEO Mustafa Suleyman has also claimed the transcription model is 55% faster and 60% cheaper than ElevenLabs' offering — figures the company has not independently benchmarked.

The release extends Microsoft's push to build its own foundation models rather than relying solely on OpenAI. Earlier MAI speech models already power Copilot and Teams, and the new ones are exposed to developers through Microsoft Foundry — a signal that the internal models are graduating from internal plumbing to sellable product.

For the industry, the message is that voice is becoming a latency battleground. As agents take calls, join meetings, and act on spoken commands, whoever transcribes fastest and cheapest holds an unglamorous but compounding advantage — and for now, that is Microsoft, at about half a dollar per hour of audio.

Comments (0)

Log in to join the discussion

Log In

No comments yet