Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

A Diffusion Model Is Now Answering Phone Calls: Inception's Mercury Voice Opens With a 320ms First-Word Latency

A Diffusion Model Is Now Answering Phone Calls: Inception's Mercury Voice Opens With a 320ms First-Word Latency

Inception Labs has made Mercury Voice, a diffusion language model built for real-time voice agents, generally available to enterprise customers. The company reports a 320 ms median time to first answer token - the only model it tested under a 500 ms conversational budget - at launch pricing of $0.20/$0.75 per million tokens. All benchmarks are company-reported.

Inception Labs has made Mercury Voice generally available to enterprise customers, positioning a diffusion language model as the reasoning engine for real-time voice agents. The model was first previewed two weeks earlier alongside Mercury 2.5, and the GA release is the company's move from research demo to paid production infrastructure for phone-based agents.

The constraint it attacks is physiological, not technical: callers perceive a pause as awkward once it stretches past roughly 500 milliseconds. That is why most voice stacks today default to small non-reasoning models - the latency budget leaves no room for a model that thinks before it speaks. Mercury Voice's bet is that diffusion decoding, which refines many token positions in parallel instead of generating text one token at a time, can compress that thinking into the same window.

The headline number is a company-reported 320 millisecond median time to first answer token, with a 750 millisecond 95th percentile, measured on a test set of real customer-service prompts at low reasoning effort. Inception says Mercury Voice was the only model in its comparison with a median under the 500 ms conversational budget, and that it was 5.9x faster than GPT-6 Luna running without reasoning. On a composite of tau3-bench Telecom, Retail and Airline, IFBench and BFCL v4, the company reports it outperformed GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash, Gemini 3.5 Flash-Lite and Qwen3.5-397B. None of these figures has been independently replicated.

On paper the model is a full agent runtime rather than a responder: 128K context with up to 50K output tokens, three reasoning-effort settings, tool calling, and an OpenAI-compatible API that drops into the LLM slot of LiveKit, Pipecat, Vapi or Retell stacks. List pricing is $0.40 per million input tokens and $1.50 per million output, halved to $0.20 and $0.75 at launch; Inception estimates a typical voice-agent workload costs about $0.009 per conversation minute - a model-layer estimate that excludes transcription, text-to-speech and telephony.

Three named customers are already paying. Audivi AI uses Mercury Voice for fully automated drive-thru ordering including order changes and upsells; Altur runs voice agents that negotiate payment plans for financial institutions; and OpenCall reports median model latency near 170 milliseconds on its production workload. Those are customer-supplied numbers from specific deployments, not benchmark results, and Inception notes they do not predict performance for other callers or stacks.

The company behind it has credible backing: co-founder and CEO Stefano Ermon is a Stanford professor whose generative modeling research underpins the diffusion approach, and Inception raised a $50 million seed round led by Menlo Ventures with participation from Mayfield, Innovation Endeavors, Microsoft's M12, Snowflake Ventures, Databricks Investment and NVIDIA's NVentures, plus angel checks from Andrew Ng and Andrej Karpathy.

The honest read is that Mercury Voice is a promising architecture with unverified numbers. Its 320 ms median measures only the model stage of a call, while a deployed phone agent must also survive transcription, speech synthesis and network delays. But the trajectory matters more than any single benchmark: diffusion LLMs have gone from text generation to search to real-time voice in under two years, each step claiming a workload where latency and intelligence must coexist. If third-party evaluations confirm the latency figures, the industry's habit of stripping reasoning out of voice stacks starts to look like a temporary compromise rather than a law of nature.

Comments (0)

Log in to join the discussion

Log In

No comments yet