In MLPerf® Training v6.0, CoreWeave trained DeepSeek-V3 671B to the benchmark’s quality target in 2.02 minutes on 8,192 NVIDIA GB300 NVL72 GPUs. It was the fastest result in the round on that benchmark.
This whitepaper is the engineering story behind that number. It walks through how CoreWeave scaled frontier training across five system configurations, from 64 GPUs to 8,192, and lays out what it took to hold strong-scaling efficiency at the largest GB300 NVL72 clusters in the round.
What’s inside:
- The full v6.0 results across DeepSeek-V3 671B, Llama 3.1 405B, Llama 3.1 8B, and GPT-OSS 20B
- A year-over-year read on Llama 3.1 405B, where time-to-train dropped 2.8x, from 27.33 minutes to 9.77 minutes
- The NCCL tuning, topology-aware scheduling, and CoreWeave Mission Control validation behind the results
- The same infrastructure customers train on today, with no benchmark-only cluster
Read this if you’re deciding where to run frontier training and you want the methodology behind the numbers.