Observability and continuous improvement for production agents
About this session
Offline evaluations can validate expected behavior, but they rarely capture the long tail of edge cases, tool failures, and evolving user interactions that emerge in production. Building reliable agents requires a continuous feedback loop that connects production telemetry, evaluation, and iteration. In this session, we'll explore how CoreWeave helps you continuously improve production agents, from instrumenting end-to-end traces to running targeted evaluations against real production data. We'll also demonstrate how coding agents such as Claude Code can leverage CoreWeave to analyze failures and continuously refine prompts, workflows, and application logic.
Share this session


Get in the room. San Francisco, September 29.
Fully Connected 2026 is where the engineers, leaders, and operators running AI in production come together for three days of depth, access, and real conversation.