> ## Content Index
> Fetch the complete content index at: https://www.notatechguy.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# CoreWeave ships NVIDIA Vera Rubin, Cognition reports 4.8x speedup
- URL: https://www.notatechguy.com/coreweave-ships-nvidia-vera-rubin-cognition-reports-4-8x-speedup/
- Published: 2026-09-30T19:20:55.000Z
- Updated: 2026-09-30T19:20:55.000Z
- Description: CoreWeave is among the first clouds to deliver NVIDIA's Vera Rubin NVL72, with Cognition reporting 4.8x higher token throughput on real engineering
- Author: Marcello Babbili
- Tags: Technology & AI, Nvidia, AI Agents, OpenAI

CoreWeave announced NVIDIA's Vera Rubin NVL72 this week, and Cognition, its first production customer, reports 4.8 times higher token throughput than the GB200 NVL72 it replaces [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com).

That 4.8x figure is Cognition's own result from early tests on one workload, with no third-party verification, no published error bars, and no released code for independent reproduction [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com).

**My read:** This is the first production deployment of Vera Rubin I've seen with a concrete benchmark attached, and 4.8x is a large enough jump to matter for anyone running agent inference at scale. But the source is an NVIDIA corporate blog quoting its own launch customer on a workload Cognition designed itself. I'd want to see a third-party reproduction before treating 4.8x as anything more than a best-case ceiling. What I find more credible is the operational detail: CoreWeave stood up a production cluster in days, and Cognition scaled to thousands of GPUs in nine months on the same platform. That deployment speed is the real signal for AI teams weighing cloud providers.

### Cognition measured the gain against GB200 on real coding tasks

CoreWeave received its first Vera Rubin NVL72 production racks earlier this month [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). Within days it built out a production cluster for Cognition [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). Cognition then ran benchmarks comparing Vera Rubin to a GB200 NVL72 baseline, drawing tasks from FrontierCode, a collection of real coding problems that AI agents were asked to solve [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). The 4.8x throughput gain was recorded on SWE-2 inference workloads — the compute spent actually generating model outputs [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com).

Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave, having scaled to thousands of GPUs in nine months [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). Silas Alberti, a founding team member at Cognition, said that having the workload on one platform with NVIDIA and CoreWeave engineers working alongside his team "matters more to us than any single spec" [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com).

That integration pitch reflects a broader truth about the inference layer, where AI economics either work or break. A 4.8x throughput jump would reshape the cost equation for any team running coding agents at volume if it holds. For a platform engineer running Devin-style coding agents on GB200 today, a genuine 4.8x throughput gain would mean either serving nearly five times the request volume on the same rack or cutting inference spend by roughly four-fifths — assuming the benchmark translates to their own workloads [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com).

![Vera Rubin NVL72 vs GB200 NVL72: token throughput (Cognition-reported, SWE-2 workload)](https://storage.ghost.io/c/6e/89/6e896869-22ef-4281-a213-b4c462c17cff/content/images/2026/09/chart_0c570e6ff83f987808ba.png)

### CoreWeave Forge and the Vera CPU expand the launch

CoreWeave also introduced CoreWeave Forge, a unified environment designed to let teams train, evaluate and refine models and agents on NVIDIA accelerated computing within a single workflow [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). The company will offer the NVIDIA Vera CPU, which NVIDIA describes as the first processor purpose-built for AI agents [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). NVIDIA reports more than 3x faster agentic sandbox startup times, though the blog offers little supporting detail, no methodology, and no independent confirmation of that figure [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com).

The Vera CPU is not yet available. CoreWeave says it "will come" to the cloud [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com).

Ian Buck, NVIDIA's vice president of hyperscale and high-performance computing, noted that CoreWeave's NVIDIA V100 GPUs are still running customer workloads nearly a decade after the Volta architecture launched [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). That longevity matters for capital decisions around AI factory build-outs.

Capacity on Vera Rubin can be operated through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference [S¹](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com). For an AI engineering team evaluating cloud providers, the practical move is to request early access and benchmark a representative task set against a current GB200 deployment rather than trusting Cognition's number on Cognition's workload.

The deep-research trail adds context. A paper on arXiv from researchers at Stanford and Cursor Research, "Mixture-of-Kittens: MoE Megakernel for NVL72s," suggests the NVL72 architecture is already attracting independent optimization work [P⁴](https://arxiv.org/html/2609.36070?ref=notatechguy.com). CoreWeave's own infrastructure tooling is partially public: its ml-containers repository on GitHub has 51 stars and an MIT licence [P²](https://github.com/coreweave/ml-containers/?ref=notatechguy.com), while its substrate repository has no stars and an unasserted licence [P³](https://github.com/coreweave/substrate?ref=notatechguy.com).

NVIDIA's infrastructure push now spans training, inference and agent execution, from cloud to automotive.

CoreWeave has not published pricing for Vera Rubin instances, and the Vera CPU has no release date beyond "will come."

---

*Sources: [S1 — From Training to Production, NVIDIA and CoreWeave Close the Loop on Ag](https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/?ref=notatechguy.com) · [P2 — coreweave/ml-containers](https://github.com/coreweave/ml-containers/?ref=notatechguy.com) · [P3 — coreweave/substrate](https://github.com/coreweave/substrate?ref=notatechguy.com) · [P4 — Mixture-of-Kittens: MoE Megakernel for NVL72s](https://arxiv.org/html/2609.36070?ref=notatechguy.com)*

---

*Written from 4 sourced items, 4 of them primary.*

## More from Not A Tech Guy

- [OpenAI's GPT-6.1 Sol targets Astra quality at one-fifth price](https://www.notatechguy.com/openai-s-gpt-6-1-sol-targets-astra-quality-at-one-fifth-price/)
- [LLM pipelines lose 40 accuracy points at interfaces](https://www.notatechguy.com/llm-pipelines-lose-40-accuracy-points-at-interfaces/)
- [Holo4 27B scores 61.7% on OSWorld 2.0, trails Opus 5.5](https://www.notatechguy.com/holo4-27b-scores-61-7-on-osworld-2-0-trails-opus-5-5/)