Inception Labs has made Mercury Voice generally available to enterprise customers, positioning a diffusion language model as the reasoning engine for real-time voice agents. The model was first previewed two weeks earlier alongside Mercury 2.5, and the GA release is the company's move from research demo to paid production infrastructure for phone-based agents.
The constraint it attacks is physiological, not technical: callers perceive a pause as awkward once it stretches past roughly 500 milliseconds. That is why most voice stacks today default to small non-reasoning models - the latency budget leaves no room for a model that thinks before it speaks. Mercury Voice's bet is that diffusion decoding, which refines many token positions in parallel instead of generating text one token at a time, can compress that thinking into the same window.
The headline number is a company-reported 320 millisecond median time to first answer token, with a 750 millisecond 95th percentile, measured on a test set of real customer-service prompts at low reasoning effort. Inception says Mercury Voice was the only model in its comparison with a median under the 500 ms conversational budget, and that it was 5.9x faster than GPT-6 Luna running without reasoning. On a composite of tau3-bench Telecom, Retail and Airline, IFBench and BFCL v4, the company reports it outperformed GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash, Gemini 3.5 Flash-Lite and Qwen3.5-397B. None of these figures has been independently replicated.
On paper the model is a full agent runtime rather than a responder: 128K context with up to 50K output tokens, three reasoning-effort settings, tool calling, and an OpenAI-compatible API that drops into the LLM slot of LiveKit, Pipecat, Vapi or Retell stacks. List pricing is $0.40 per million input tokens and $1.50 per million output, halved to $0.20 and $0.75 at launch; Inception estimates a typical voice-agent workload costs about $0.009 per conversation minute - a model-layer estimate that excludes transcription, text-to-speech and telephony.
Three named customers are already paying. Audivi AI uses Mercury Voice for fully automated drive-thru ordering including order changes and upsells; Altur runs voice agents that negotiate payment plans for financial institutions; and OpenCall reports median model latency near 170 milliseconds on its production workload. Those are customer-supplied numbers from specific deployments, not benchmark results, and Inception notes they do not predict performance for other callers or stacks.
The company behind it has credible backing: co-founder and CEO Stefano Ermon is a Stanford professor whose generative modeling research underpins the diffusion approach, and Inception raised a $50 million seed round led by Menlo Ventures with participation from Mayfield, Innovation Endeavors, Microsoft's M12, Snowflake Ventures, Databricks Investment and NVIDIA's NVentures, plus angel checks from Andrew Ng and Andrej Karpathy.
The honest read is that Mercury Voice is a promising architecture with unverified numbers. Its 320 ms median measures only the model stage of a call, while a deployed phone agent must also survive transcription, speech synthesis and network delays. But the trajectory matters more than any single benchmark: diffusion LLMs have gone from text generation to search to real-time voice in under two years, each step claiming a workload where latency and intelligence must coexist. If third-party evaluations confirm the latency figures, the industry's habit of stripping reasoning out of voice stacks starts to look like a temporary compromise rather than a law of nature.
Comments (0)
Log in to join the discussion
Log InNo comments yet