INTELLIGENT AGENT OBSERVABILITY

CoreWeave Agent Lens

Intelligent agent observability at enterprise scale. Agent Lens turns tens of millions of traces into automated actions that continuously improve production agent performance.

Automate agent improvement from production traces to proven fixes

Production agents fail in ways that are hard to catch and even harder to trace, from drifted outputs and broken tool calls to prompt changes that silently cause regressions. Agent Lens brings the entire loop together, from detecting failures to validating and shipping fixes, so your organization works from the same trace across prompts, harnesses, and models without stitching together disparate tools.

Automate agent improvement

When an agent fails in production, tracing it back usually means bouncing between different tools and losing the thread at every handoff. Agent Lens powers the entire agent improvement loop, from tracing to running evals, so nothing gets lost in the handoff and nothing occurs unnoticed.

Insights for every role

Agent Lens gives every role the tools and insights they need. Product managers and domain experts can review conversations, define quality, create evaluations, and monitor performance, while AI engineers can dive into traces, tool calls, latency, and other technical details to diagnose and fix issues. Everyone works from the same context, with the depth and workflows they need.

Built for enterprise

From healthcare and financial services to technology and AI-native enterprises, every industry has different requirements for security, governance, privacy, and infrastructure. Agent Lens is designed to accommodate them, with enterprise controls such as PII redaction and flexible deployment options.

OBSERVE

Understand every trace at a glance

Follow agent behavior as a conversation instead of decoding raw JSON or nested spans. Agent Lens renders every turn in plain dialogue, giving engineers, product teams, and subject-matter experts a clear view of what the agent said and did.

When you need to go deeper, expand any turn to inspect every input, output, tool call, model response, and piece of metadata behind it, so technical depth is always available without sacrificing readability.

DETECT

Catch issues before they become patterns

Agent Lens analyzes tens of millions of production traces to automatically uncover, group, and prioritize failures based on frequency and severity. Out of the box, it turns massive volumes of agent behavior into actionable insights, helping teams quickly understand which issues matter most. Optimized for CoreWeave inference, Agent Lens can continuously analyze production traffic at scale to detect emerging failures and behavioral drift. When something goes wrong, it connects the failure to the underlying evidence and root cause, so teams can move from detection to diagnosis and fix without exporting logs or switching tools.

EVALUATE

Prove it before you ship

A fix that feels right isn’t the same as a fix that has been verified. From any failure cluster, you can create a custom LLM judge and tune it to your standards. Domain experts define what should have happened, review discrepancies between human and judge scores, and improve alignment before the judge automatically evaluates every new production trace.

Each detected failure becomes a test case in a growing evaluation set. Run a candidate fix against the entire set, compare results side by side, and catch regressions before promoting the change to production. You ship because you tested the fix against real cases, not because you hope it works.

Signals running on CoreWeave Inference are preconfigured and optimized out of the box. Connect through OpenTelemetry or a native integration, and automated evaluation begins within minutes, without a lengthy setup project.

IMPROVE

Turn insights into automated improvement

Agent Lens analyzes production conversations to identify common user intents and capability gaps. It clusters related requests by frequency and shows how they trend over time, helping product managers and engineers prioritize improvements based on real customer evidence.

Connect coding agents like Claude Code to Agent Lens through Skills and MCP. They can work with live production data, run evaluations, test changes, and execute automated improvement loops—turning what you learn in production into action without leaving the workflow.

work with weave

Agent Lens and Weave

Agent Lens complements W&B Weave. If you use Weave to observe production agents and evaluate changes, you can use Agent Lens at no additional cost to analyze production traces faster, cluster related failures and user intents, and generate actionable insights. These insights strengthen continuous agent improvement by helping teams prioritize issues, identify capability gaps, and turn real-world evidence into verified improvements.

ENTERPRISE READY

Built for enterprise scale and control

Deploy Agent Lens where your environment requires, keep sensitive data where it belongs, and connect production insights back to model development.

Deploy on your terms

Run Agent Lens as SaaS, in a dedicated cloud, or on-premises—so deployment fits your infrastructure and security requirements.

Keep control of your data

Bring your own bucket to keep regulated or security-sensitive data where it needs to live.

Close the loop back to training

Connect Agent Lens with Weights & Biases Models so production failures become signals for future training runs—not issues stranded in a bug tracker.

Scale with lower cost

With CoreWeave’s cost-efficient inference, Agent Lens optimized the performance teams need to continuously improve agents as usage grows.

AGENT LENS IS PART OF COREWEAVE FORGE

Connected to the entire AI development lifecycle

A recurring failure is a training signal, not just a bug. Because Agent Lens connects to Weights & Biases Models, what your team learns from production agents can drive the next training run — connecting agent behavior to the full AI development lifecycle on CoreWeave.

Run

Observe

Curate

Improve

Evaluate

FAQS

Frequently asked questions

What is CoreWeave Agent Lens?

How does Agent Lens help improve AI agents?

How does Agent Lens identify problems in production agents?

Can non-engineers use Agent Lens to review agent behavior?

How does Agent Lens connect production insights back to model development?

Related resources

The same news reads differently depending on where you sit. Here’s the version that applies to you.

BLOG

Automate Agent Observability and Improvement with CoreWeave Agent Lens

VIDEO

Agent Lens Explainer Video

VIDEO

Forge explainer video

Builder Resource Center: Learn from every run

Explore demos, code, and technical resources for every stage of the AI loop. Learn how researchers, developers, and CoreWeave engineers build, observe, evaluate, and improve AI systems—and put those insights to work.

GET STARTED

Get started with Agent Lens

Intelligent agent observability and improvement built for enterprise scale. See exactly what your agents do in production, discover failure patterns you didn't know to look for, and turn every failure into a test that stops it recurring.