Celebrating the One Billion Runs Milestone on Weights & Biases

How the AI community helped us reach an incredible milestone, and what it means for the next generation of AI breakthroughs
Celebrating the One Billion Runs Milestone on Weights & Biases

When we started Weights & Biases in 2017, we had a single goal in mind: to build the best tools for machine learning practitioners. Our first tool was an experiment tracker we built in a karate studio in San Francisco and teams like OpenAI, Toyota Research, and Uber picked it up quickly, replacing TensorBoard or internal solutions that weren’t quite up to the task. And though AI progress has accelerated at an incredible pace in the decade since we launched, we’ve kept building the tools our customers need to continue the research that continues transforming our society. 

We’re incredibly proud of the work we’ve done and how far we’ve come. And while we spend every day looking forward to designing the next indispensable tool, today, we want to look back. That’s because we just hit a milestone that was almost unthinkable in 2017: more than a billion runs have been tracked in Weights & Biases. 

First off, we want to say “thank you” to everyone who’s ever used Weights & Biases. Whether you were a student learning how to build your first computer vision model on MNIST or a researcher building the latest foundation model, we’re honored you chose us to help you on your journey. Our focus over the past nine years has never changed and it won’t going forward. We’ll keep building the best tools we can for AI teams of every size. And while I’ll talk a bit about what’s next, I really wanted to take a moment to express our gratitude. This is an incredible milestone.

But it’s only the beginning. 

The road to a trillion runs

The size and complexity of AI models is constantly increasing. Right now, a single foundation-scale effort can log hundreds of thousands of experiments in a year, a scale that would have been genuinely impossible in 2017. But as labs have started applying AI to their own training, we think the experiment load will increase massively along with it. 

We’re explicitly building for this next era. Outside of the performance work we always prioritize, we recently launched CoreWeave ARIA, which stands for AI Research and Iteration Agent. ARIA lives in Weights & Biases and helps analyze experiments, generate hypotheses, launch new runs, and evaluate results that can automate the AI research loop. 

ARIA is just one tool that supports the new reality in our space. When we started Weights & Biases, the AI model development loop was largely the purview of machine learning researchers. That’s changed. It’s far more common now that engineers ship agents, domain experts review outputs, internal teams build and tweak agents, and everything has to be meticulously maintained, evaluated, and kept performant. Simply put, that means far more people are involved in the process. Building the tools for both researchers and the myriad other stakeholders is a challenge we’re excited to take on. 

Joining CoreWeave has been a godsend for that mission. Now, the tools you’re building and the infra you build on work as a single system: training and inference, tracking and evaluation, observability and feedback, all the way down to the hardware. A unified platform built for the entire loop was why we brought our companies together. I’m thrilled with how it’s gone.

The stories behind a billion runs

All that said: a billion is a big number but it’s just a number. We think the stories behind it are far more interesting. Here are just a few:

Decart developed Oasis 3, the first API accessible world model for physical AI on Weights & Biases. They centralized experiment tracking, hyperparameter optimization, dataset and model versioning, and lineage tracking through W&B Models so the team could accelerate large scale training, reproduce results with confidence, and deliver production ready world models for robotics and autonomous systems.

Pinterest standardized machine learning development across their organization with Weights & Biases as the foundation for experiment tracking and model management. Integrated into its MLEnv framework, Weights & Biases enables hundreds of ML engineers to track experiments, version model artifacts with W&B Registry, manage fault tolerant training checkpoints, and accelerate iteration across hundreds of thousands of training jobs each month.

Scaled Cognition trains highly reliable LLMs for regulated industries like banking and healthcare on Weights & Biases. With low friction SDK integration, comprehensive experiment tracking, live system monitoring, and reproducible run histories, Weights & Biases helps the team accelerate model development, optimize custom training algorithms, and deliver AI models that achieve 114% higher accuracy than leading general purpose LLMs for enterprise customer support.

MasterClass uses W&B Weave to evaluate and improve AI teaching agents that deliver personalized instruction and real time feedback. By tracing complete conversations, monitoring production behavior, and using signals and custom monitors to evaluate live learner interactions, W&B Weave helps the team continuously improve teaching quality and scale personalized learning with greater confidence.

What’s next

In 2017, transformers were a fresh idea. Dall-E was four years away. ChatGPT another year after that. Which is all to say, AI has changed a ton since we founded Weights & Biases in that karate studio. And while the tools we need will change with every new research breakthrough, our fundamental mission never will: we’ll always be committed to making the best tools for AI. 

While not every experiment bears fruit, every run is a step towards better models. But even after a billion, there’s so much left to build together. I’m excited to see where we go next. I hope you’ll be there with us. 

Join us at Fully Connected 2026, from September 29 - October 1 at Moscone South in San Francisco, to help build what’s next. 

Celebrating the One Billion Runs Milestone on Weights & Biases

From 1 billion runs to autonomous AI research, explore the latest Weights & Biases innovations for AI model and agent development.

Related Blogs

CoreWeave Cloud,
Copy code
Copied!