Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

CoreWeave Puts NVIDIA's Vera Rubin NVL72 Into Production, With Cognition as First Customer Reporting Up to 4.8x Token Throughput

CoreWeave Puts NVIDIA's Vera Rubin NVL72 Into Production, With Cognition as First Customer Reporting Up to 4.8x Token Throughput

NVIDIA says AI cloud provider CoreWeave has moved Vera Rubin NVL72 into production, with AI software firm Cognition as the first customer running production workloads on it. Cognition's early tests showed up to 4.8x total token throughput versus GB200 NVL72, and CoreWeave reported agent sandboxes starting over 3x faster on Vera CPUs. All figures are vendor-reported.

The Rubin generation has moved from roadmap to revenue. NVIDIA said this week that CoreWeave, the AI cloud operator and its most aggressive early-adopter, has placed Vera Rubin NVL72 systems into production and is selling capacity on them to customers — with AI software firm Cognition becoming the first company to run production workloads on the platform, according to a statement carried by Chinese financial wire Cailianshe on September 30.

The headline number comes from Cognition, the company behind the Devin coding agent. Its early testing found that Vera Rubin NVL72 delivered up to 4.8 times the total token throughput of the previous-generation GB200 NVL72 in real software-engineering AI inference tasks — exactly the long-horizon, tool-calling workloads that agentic coding products hammer all day. NVIDIA's statement frames it as an early result; independent numbers will take longer.

CoreWeave contributed a second datapoint: agent sandboxes running on NVIDIA's Vera CPU start up more than 3 times faster than before. That matters for a different reason than raw model speed. Agentic platforms spin up isolated environments thousands of times a day, and sandbox startup latency has quietly become one of the hidden costs of the agent era — every second of boot time is money and user-visible delay.

Vera Rubin pairs NVIDIA's next-generation Rubin GPUs with its purpose-built Vera CPU, the chip NVIDIA has been positioning as the control plane for agentic AI — the same silicon that underpins its OpenShell agent-security runtime. Production deployment on CoreWeave suggests the platform has moved past pilot evaluations into workloads customers are actually paying for, ahead of the wider OEM server wave.

The throughput claim, if it holds up under third-party measurement, would be the latest step in a escalating benchmark war defined less by FLOPS than by token economics. Cognition's business depends on it: coding agents bill by outcome but pay for tokens, so a 4.8x throughput gain translates almost linearly into gross margin. That is why the first production customer on new hardware is now a marketing slot that cloud providers compete to fill.

Worth keeping the vendor label on all of it: the 4.8x and 3x figures come from the companies selling and buying the hardware, tested on their own workloads, with no published methodology. But the deployment itself is verifiable and notable — Rubin-class systems running paying production traffic within weeks of launch is a signal about how compressed the AI hardware cycle has become.

For the broader market, the message to enterprises still standardizing on Blackwell is uncomfortable: the generation they are buying now already has a successor earning revenue elsewhere. The token-throughput arms race is resetting annually, and procurement cycles have not caught up.

Comments (0)

Log in to join the discussion

Log In

No comments yet