The Rubin generation has moved from roadmap to revenue. NVIDIA said this week that CoreWeave, the AI cloud operator and its most aggressive early-adopter, has placed Vera Rubin NVL72 systems into production and is selling capacity on them to customers — with AI software firm Cognition becoming the first company to run production workloads on the platform, according to a statement carried by Chinese financial wire Cailianshe on September 30.
The headline number comes from Cognition, the company behind the Devin coding agent. Its early testing found that Vera Rubin NVL72 delivered up to 4.8 times the total token throughput of the previous-generation GB200 NVL72 in real software-engineering AI inference tasks — exactly the long-horizon, tool-calling workloads that agentic coding products hammer all day. NVIDIA's statement frames it as an early result; independent numbers will take longer.
CoreWeave contributed a second datapoint: agent sandboxes running on NVIDIA's Vera CPU start up more than 3 times faster than before. That matters for a different reason than raw model speed. Agentic platforms spin up isolated environments thousands of times a day, and sandbox startup latency has quietly become one of the hidden costs of the agent era — every second of boot time is money and user-visible delay.
Vera Rubin pairs NVIDIA's next-generation Rubin GPUs with its purpose-built Vera CPU, the chip NVIDIA has been positioning as the control plane for agentic AI — the same silicon that underpins its OpenShell agent-security runtime. Production deployment on CoreWeave suggests the platform has moved past pilot evaluations into workloads customers are actually paying for, ahead of the wider OEM server wave.
The throughput claim, if it holds up under third-party measurement, would be the latest step in a escalating benchmark war defined less by FLOPS than by token economics. Cognition's business depends on it: coding agents bill by outcome but pay for tokens, so a 4.8x throughput gain translates almost linearly into gross margin. That is why the first production customer on new hardware is now a marketing slot that cloud providers compete to fill.
Worth keeping the vendor label on all of it: the 4.8x and 3x figures come from the companies selling and buying the hardware, tested on their own workloads, with no published methodology. But the deployment itself is verifiable and notable — Rubin-class systems running paying production traffic within weeks of launch is a signal about how compressed the AI hardware cycle has become.
For the broader market, the message to enterprises still standardizing on Blackwell is uncomfortable: the generation they are buying now already has a successor earning revenue elsewhere. The token-throughput arms race is resetting annually, and procurement cycles have not caught up.
Comments (0)
Log in to join the discussion
Log InNo comments yet