CoreWeave announced NVIDIA's Vera Rubin NVL72 this week, and Cognition, its first production customer, reports 4.8 times higher token throughput than the GB200 NVL72 it replaces S¹.
That 4.8x figure is Cognition's own result from early tests on one workload, with no third-party verification, no published error bars, and no released code for independent reproduction S¹.
My read: This is the first production deployment of Vera Rubin I've seen with a concrete benchmark attached, and 4.8x is a large enough jump to matter for anyone running agent inference at scale. But the source is an NVIDIA corporate blog quoting its own launch customer on a workload Cognition designed itself. I'd want to see a third-party reproduction before treating 4.8x as anything more than a best-case ceiling. What I find more credible is the operational detail: CoreWeave stood up a production cluster in days, and Cognition scaled to thousands of GPUs in nine months on the same platform. That deployment speed is the real signal for AI teams weighing cloud providers.
Cognition measured the gain against GB200 on real coding tasks
CoreWeave received its first Vera Rubin NVL72 production racks earlier this month S¹. Within days it built out a production cluster for Cognition S¹. Cognition then ran benchmarks comparing Vera Rubin to a GB200 NVL72 baseline, drawing tasks from FrontierCode, a collection of real coding problems that AI agents were asked to solve S¹. The 4.8x throughput gain was recorded on SWE-2 inference workloads — the compute spent actually generating model outputs S¹.
Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave, having scaled to thousands of GPUs in nine months S¹. Silas Alberti, a founding team member at Cognition, said that having the workload on one platform with NVIDIA and CoreWeave engineers working alongside his team "matters more to us than any single spec" S¹.
That integration pitch reflects a broader truth about the inference layer, where AI economics either work or break. A 4.8x throughput jump would reshape the cost equation for any team running coding agents at volume if it holds. For a platform engineer running Devin-style coding agents on GB200 today, a genuine 4.8x throughput gain would mean either serving nearly five times the request volume on the same rack or cutting inference spend by roughly four-fifths — assuming the benchmark translates to their own workloads S¹.

CoreWeave Forge and the Vera CPU expand the launch
CoreWeave also introduced CoreWeave Forge, a unified environment designed to let teams train, evaluate and refine models and agents on NVIDIA accelerated computing within a single workflow S¹. The company will offer the NVIDIA Vera CPU, which NVIDIA describes as the first processor purpose-built for AI agents S¹. NVIDIA reports more than 3x faster agentic sandbox startup times, though the blog offers little supporting detail, no methodology, and no independent confirmation of that figure S¹.
The Vera CPU is not yet available. CoreWeave says it "will come" to the cloud S¹.
Ian Buck, NVIDIA's vice president of hyperscale and high-performance computing, noted that CoreWeave's NVIDIA V100 GPUs are still running customer workloads nearly a decade after the Volta architecture launched S¹. That longevity matters for capital decisions around AI factory build-outs.
Capacity on Vera Rubin can be operated through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference S¹. For an AI engineering team evaluating cloud providers, the practical move is to request early access and benchmark a representative task set against a current GB200 deployment rather than trusting Cognition's number on Cognition's workload.
The deep-research trail adds context. A paper on arXiv from researchers at Stanford and Cursor Research, "Mixture-of-Kittens: MoE Megakernel for NVL72s," suggests the NVL72 architecture is already attracting independent optimization work P⁴. CoreWeave's own infrastructure tooling is partially public: its ml-containers repository on GitHub has 51 stars and an MIT licence P², while its substrate repository has no stars and an unasserted licence P³.
NVIDIA's infrastructure push now spans training, inference and agent execution, from cloud to automotive.
CoreWeave has not published pricing for Vera Rubin instances, and the Vera CPU has no release date beyond "will come."
Sources: S1 — From Training to Production, NVIDIA and CoreWeave Close the Loop on Ag · P2 — coreweave/ml-containers · P3 — coreweave/substrate · P4 — Mixture-of-Kittens: MoE Megakernel for NVL72s
Written from 4 sourced items, 4 of them primary.