GPU choice that matches the workload
Pick the GPU class that fits your latency, throughput, and cost targets. Single-node or distributed multi-node serving, billed per GPU-hour.
Bring your own weights, stored in CoreWeave Object Storage
Point deployments at fine-tuned checkpoints, custom architectures, or OSS weights in CoreWeave Object Storage.
Open runtimes, OpenAI-compatible endpoints
vLLM and SGLang runtimes with OpenAI-compatible endpoints out of the box. Swap models, runtimes, or GPU classes without rebuilding the serving stack.
Gateway-managed routing and traffic control
A tenant-isolated gateway runs authentication, load balancing, and request routing across replicas. Optimize for latency, data locality, or compliance.
Cost that maps to infrastructure
Per GPU-hour billing against your chosen GPU class and capacity model. No egress fees, no ingress fees, no service markup.





