CoreWeave Leads Cloud Providers in MLPerf® Inference v6.1 Performance with NVIDIA Blackwell Ultra

From multimodal to MoE reasoning models, CoreWeave delivers leading performance across today’s most demanding inference workloads
CoreWeave Leads Cloud Providers in MLPerf® Inference v6.1 Performance with NVIDIA Blackwell Ultra

As AI models evolve, inference throughput must be evaluated and validated in context. A token generated by a text-only model does not place the same demands on infrastructure as one generated by a multimodal or MoE reasoning model. Multimodal architectures add compute-intensive image encoding and cross-attention, while MoE reasoning models route tokens through specialized experts and often generate longer responses. These differences change the compute, memory, and networking required to sustain a given tokens-per-second rate and ultimately what that throughput costs. 

With MLPerf® Inference v6.1, throughput is evaluated across a diverse set of model architectures. CoreWeave participated in the Datacenter Closed division’s Available category, submitting results across four NVIDIA platforms and four model families that are in use today.  NVIDIA Blackwell and Blackwell Ultra are commonly adopted platforms in production today. The models benchmarked are the following:

  • Qwen3-VL-235B-A22B, multimodal AI powering production search and catalog use cases
  • DeepSeek-R1-671B, one of the largest MoE reasoning models and widely adopted
  • GPT-OSS-120B, the open-weight reasoning model seeing rapid adoption since release
  • Llama 2 70B, one of the most widely deployed models in production

Together, these throughput numbers on the infrastructure and models are already doing the work in production and delivering real world performance.

What the results mean for customers

Benchmarks capture performance at a moment in time, but the capabilities behind that performance continue to compound. CoreWeave extracts more value from deployed hardware while rapidly bringing the latest silicon and emerging model classes to run at production-scale.

For customers evaluating inference infrastructure, two measures matter most: tokens per GPU1 and time to production. Higher per-GPU throughput can reduce the infrastructure required to serve a workload, while rapid access to new platforms helps teams move from model development to production sooner.

CoreWeave’s performance across multimodal, frontier-scale reasoning, MoE, and large language models gives customers the efficiency and scale to turn NVIDIA accelerated computing into real world AI work.

CoreWeave delivers the highest Qwen3-VL-235B-A22B server throughput among cloud providers

CoreWeave’s results with Qwen3-VL-235B-A22B highlight how our software stack turns new multimodal scenarios and rack-scale infrastructure into leading inference performance.  CoreWeave delivered the highest server throughput among cloud providers using an NVIDIA GB300 NVL72 rack. In server scenario, CoreWeave sustained 1,196 queries per second, the highest server throughput among cloud providers, and 1,135 samples per second in offline scenario. 

CoreWeave increases DeepSeek-R1-671B per GPU throughput by nearly 20% over MLPerf® v6.0

Continuous software optimization increases the throughput of GPUs already in production, helping customers serve more tokens and improve inference economics without waiting for the next hardware generation. Significant gains can come from regular tuning that delivers more throughput which is why CoreWeave continuously optimizes its serving stack, scheduling, and operations to unlock more performance from existing infrastructure. The results are measurable.  Just five months after MLPerf® Inference v6.0, comparing the 72-GPU v6.1 submission with the 64-GPU v6.0 submission on NVIDIA GB200 NVL72,derived per-GPU Server throughput increased by 19.8%. 

CoreWeave achieves the highest GPT-OSS-120B per GPU throughput  

On a single NVIDIA GB300 NVL72 rack, CoreWeave sustained 16,635 tokens per second per GPU in offline scenario, and 16,118 in server scenario. These were the highest per-GPU throughput of any v6.1 Datacenter Closed submission on GPT-OSS-120B, on any silicon, in both scenarios.1 CoreWeave achieved over 1.16 million tokens per second in server and over 1.19 million tokens per second in offline scenarios on a single NVIDIA GB300 NVL72 rack.

Per-GPU throughput maps directly to inference economics: At a given GPU price, serving more tokens per accelerator can reduce infrastructure cost per token. CoreWeave’s advantage also extends across platforms: Our NVIDIA GB200 NVL72 posted the highest rack-scale GPT-OSS-120B total in its class, reaching 901,058 tokens per second in server and 911,566 tokens per second in offline scenarios. Our NVIDIA HGX B200 and HGX B300 systems led all HGX submissions from cloud providers in both scenarios.

CoreWeave delivers the highest throughput among cloud providers with Llama 2 70B 

With Llama 2 70B, CoreWeave demonstrated rack-scale serving capacity for a workhorse model that is widely deployed. We delivered the highest throughput among cloud providers with NVIDIA GB300 NVL72 with a throughput of 944,902 tokens per second in server scenario and 1,136,100 in offline scenario. In the offline scenario, we achieved the only result above one million tokens per second among cloud provider submissions using NVIDIA GB300 NVL72.    

Summary of CoreWeave’s Performance Results 

CoreWeave’s MLPerf® Inference v6.1 submissions covered four distinct model families with multiple NVIDIA Blackwell and Blackwell Ultra platforms:

Model Workload NVIDIA platform CoreWeave Results
Qwen3-VL-235B-A22B Multimodal AI GB300 NVL72 Highest server throughput among cloud providers with NVIDIA GB300 NVL72.
DeepSeek-R1 671B MoE reasoning HGX B200 and GB200 NVL72 Increased per-GPU server throughput by nearly 20% over our MLPerf® v6.0 result on NVIDIA GB200 NVL72.
GPT-OSS-120B Reasoning HGX B200, HGX B300, GB200 NVL72, GB300 NVL72 Highest per-GPU throughput of any submission, server and offline1; highest aggregate throughput of any NVIDIA-based system, on a single GB300 NVL72 rack
Llama 2 70B Large language model HGX B200, HGX B300, GB200 NVL72, GB300 NVL72 Highest total server and offline throughput with NVIDIA GB300 NVL72 among cloud providers

Benchmarks were performed on production clusters

These results were not produced on a tuned cluster optimized just for benchmarking. CoreWeave ran its submission on the same production clusters, images, and operational stack available to customers:

  • CoreWeave Mission Control provides fleet-wide health monitoring and lifecycle management, with continuous benchmarking that ties system behavior to sustained tokens-per-second outcomes.
  • Topology-aware scheduling via SUNK pins inference workloads inside high-bandwidth NVIDIA NVLink domains on NVIDIA GB200 NVL72 and GB300 NVL72 systems.

The same discipline that lets CoreWeave stand up four Blackwell platforms—NVIDIA HGX B200, HGX B300, GB200 NVL72 and GB300 NVL72 in one submission window is what moves customers from delivery to serving tokens in days.

To learn more: 

Join us at Fully Connected 

Watch Many Workloads, One Cluster: Training and Inference Without Idle GPUs

Read the Signal65 report on What does AI infrastructure really cost?    


Results referenced: MLCommons® MLPerf® Inference v6.1, Datacenter Closed division, available category, published September 16, 2026. See www.mlcommons.org for complete results.

1 Per-GPU throughput is a derived metric calculated as total system throughput divided by the number of accelerators. Result not verified by MLCommons Association. The MLPerf® name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited.

CoreWeave Leads Cloud Providers in MLPerf® Inference v6.1 Performance with NVIDIA Blackwell Ultra

See how CoreWeave’s full-stack optimizations delivered leading inference throughput for NVIDIA Blackwell and Blackwell Ultra platforms

Related Blogs

Copy code
Copied!