NVIDIA Vera Rubin NVL72 on CoreWeave: 10x More Tokens Per Megawatt Than Blackwell

NVIDIA Vera Rubin NVL72 on CoreWeave: 10x More Tokens Per Megawatt Than Blackwell

In early June, CoreWeave completed the industry's first bring-up and validation of NVIDIA Vera Rubin NVL72, standing up the rack-scale system end-to-end, with power, cooling, networking, and compute all verified and running. That milestone showed the platform running at production scale. The next question everyone has asked: How does it perform? Today, we’re happy to share the first-ever measured silicon performance on Vera Rubin NVL72.

DeepSeek R1, one of the reasoning models driving the current surge in agentic AI, was run across both Vera Rubin NVL72 and NVIDIA Blackwell NVL72. In this evaluation, token throughput per megawatt was measured against interactivity (as TPS/user), which indicates the responsiveness a user or agent experiences while a model reasons through a task. At a matched interactivity target, Vera Rubin NVL72 generated 10x tokens-per-second per megawatt compared to NVIDIA GB200 NVL72 on the same DeepSeek R1 workload.

Line chart comparing DeepSeek R1 inference throughput per GPU against interactivity for Vera Rubin NVL72 and GB200 NVL72. The Vera Rubin NVL72 curve stays above GB200 NVL72 throughout, with the gap peaking near 10 times in the middle of the range.
While delivering similar per-user interactivity, Vera Rubin NVL72 delivers up to 10x higher tokens per megawatt per unit of energy on Vera Rubin NVL72 at similar TPS/user compared to GB200 NVL72.

What this performance improvement really means

This evaluation compares the performance on each platform with all major, state-of-the-art inference optimizations turned on. These include large-scale expert parallelism, NVFP4 precision, multi-token prediction, and disaggregated prefill and decode, all enabled using NVIDIA TensorRT-LLM and NVIDIA Dynamo.

At similar interactivity targets, Vera Rubin NVL72 delivers an order-of-magnitude improvement in token throughput per megawatt, empowering CoreWeave customers to run far more reasoning-model traffic within the same power budget or run the same traffic on far less power. That means lower cost-per-token and more room to scale at the same infrastructure investment, which is beneficial for agentic workloads that call tools and reason across many steps. 

Here's what this will look like in practice:

  • A global cybersecurity company expects to run threat detection inference at 10x better performance per watt, letting them scale agentic security across more endpoints without growing their data center footprint.
  • An autonomous coding agent company expects a similar jump in inference per GPU at a fraction of the token cost, letting agents stay on hard problems longer without breaking unit economics.
  • An AI-native search engine already running on GB200 NVL72 expects a similar jump in inference throughput per GPU, letting them expand real-time, cited-answer search to more users without breaking response-time targets.

Scaling agentic AI with NVIDIA Vera Rubin NVL72

Reasoning models like DeepSeek R1 don't answer in a single pass. They plan, check their own work, and revise, generating far more tokens per query than a typical response. Every one of those tokens has to move through memory and across the NVLink domain, which makes interactivity, not raw FLOPs alone, the number that decides whether an agent feels instant or sluggish. 

That's exactly the constraint NVIDIA Vera Rubin NVL72 was built to address. 72 NVIDIA Rubin GPUs and 36 NVIDIA Vera CPUs share a single rack connected by a 260 TB/s all-to-all NVLink 6 fabric, with native NVFP4 support built into the Transformer Engine.

We've invested heavily in the operational layer underneath it to make sure that these efficiency gains hold up once a customer's workload is running in production at scale. We've written before about how CoreWeave operates rack-scale AI as a cloud service, from the patent-pending liquid cooling in Valvey to the unified rack control in Racky, all coordinated by our Rack LifeCycle Controller.

A foundation for even more performance gains

We've seen this pattern before. When CoreWeave brought GB200 NVL72 online, initial inference numbers improved substantially over the following months as we tuned scheduling, networking, and the software stack around the hardware. This work was reflected in results like our record-breaking MLPerf inference submissions and our two-minute DeepSeek-V3 training benchmark. We expect the same trajectory here: this measurement is a starting point on Vera Rubin NVL72,  and we'll continue to publish updated numbers as optimization work progresses. 

Ready to start planning your Vera Rubin NVL72 roadmap? Contact us here.

Want to go deeper? Here’s where to look next:

NVIDIA Vera Rubin NVL72 on CoreWeave: 10x More Tokens Per Megawatt Than Blackwell

CoreWeave publishes first-ever measured silicon performance of NVIDIA Vera Rubin NVL72, marking a new milestone in AI infrastructure.

Related Blogs

GPU Compute,
Copy code
Copied!