Closed-source frontier models still lead on long, open-ended reasoning tasks. But most production agents perform a single task under a fixed set of rules, and open models now handle that well. That makes accuracy per dollar the real comparison. In this demo, we move a text-to-SQL agent from Claude Opus 4.5 to GLM-5.2 running on CoreWeave serverless inference. Then we test how six prompt engineering techniques impact accuracy and cost.
Getting started takes one change: point the OpenAI client you already use at the CoreWeave base URL and sign in with your Weights & Biases API key. You can pick a model ID from the catalog, or try the model first in the playground. There, you can adjust key parameters, bring your own tools, and compare models side by side in separate chat windows. Every run is logged to W&B Weave, so you can review each agent run afterward.
Try the notebook yourself in CoreWeave ARENA. Swap models, test your own workflow, and check it under production-like conditions before you commit.
