Serverless Inference

Call an API and start building. No clusters, no GPU selection, no infrastructure to stand up.

Why serverless inference

Start building without standing up infrastructure

With Serverless Inference, you call an API and get instant access to leading open-source models and your own fine-tuned checkpoints. CoreWeave hosts the models and runs the GPUs, so the only thing you manage is the application.

Leading open-source models or your own LoRA weights, one API

Access a curated catalog of open-source models, or bring your own LoRA weights and serve them side by side. Switch between them without changing your integration.

Tracing, evals, and observability

Every request can be traced and evaluated through native observability features, so teams can debug and improve AI applications and agents without extra instrumentation.

Playground access with zero configuration

Explore and compare models directly in the playground before writing a line of code. No endpoints to configure and no keys to manage.

Who serverless inference is built for

Move fast without standing up infrastructure

The teams that get the most value from Serverless Inference are testing, iterating, and shipping AI applications faster than they can provision infrastructure for them.

AI engineers

Iterating on AI applications and agents

Teams building copilots, agents, and RAG applications who need to test and swap models quickly without provisioning infrastructure for every experiment.

AI teams

Bringing fine-tuned models to production fast

Teams that have fine-tuned a model using LoRA and need to serve it immediately, without building a dedicated deployment pipeline first.

Platform Teams

Evaluating a workload before committing to scale

ML platform teams who need to grow into distributed inference while preserving GPU choice, runtime flexibility, and per-GPU-hour cost visibility as workloads expand.

Left
Right
CoreWeave inference paths

Inference on your terms

Three inference paths built on the award-winning CoreWeave Cloud. Move between them as your workloads evolve—without replatforming—so you consistently get predictable performance and infrastructure-aligned economics.

Serverless inference

Pay-per-token inference on a curated OSS model catalog. No clusters to manage. Built-in tracing, evals, and observability for AI applications and agents.

Best for
Rapid iteration and AI app development
infrastructure management
Low — API only
Model support
Curated OSS + LoRAs
Pricing
Pay-per-token

Dedicated inference

Deploy custom weights on explicitly chosen GPUs. CoreWeave operates routing, scaling, and lifecycle. Full execution transparency, no cluster management.

Best for
Custom model serving at production scale
infrastructure management
Medium — GPU, Zone, Runtime
Model support
Open-source or custom weights
Pricing
Pay-per-GPU-hour

Inference on CKS

Full self-managed inference on CoreWeave Kubernetes Service. Own the entire serving stack—runtimes, scheduling, autoscaling, multi-node topology—on dedicated bare-metal GPU nodes.

Best for
Full infrastructure ownership and deep tuning
infrastructure management
High — full Kubernetes control
Model support
Any
Pricing
Pay-per-GPU-hour (Reserved, On-demand, Spot, Flex)
How it helps

What can you do with Serverless Inference?

Go from open-source model to live endpoint in minutes

Call the API or use the playground. No signing up with a separate hosting provider, no cluster to provision, no GPU to select. Point your application at the endpoint and start sending requests.

Iterate on fine-tuned models without managing serving infrastructure

Bring your own LoRA weights and serve fine-tuned models without building a pipeline to fetch, hot-swap, and scale weights for every iteration. CoreWeave handles that complexity so teams can move between training and inference quickly.

See how your application behaves in production, not just at the prompt

Catch the request that failed in production, not just the one that failed at the prompt. See exactly where an agent's reasoning broke down, days after it shipped.

Left
Right
Play video
How serverless inference works

Four steps to your first API call

Choose a model, call the API, and start iterating. There is no cluster to provision and no GPU to select.

  1. Choose a model
    Pick from the curated catalog of open-source models, or point to your own LoRA weights.
  2. Call the API
    Send requests to the OpenAI-compatible endpoint using your existing SDK. No new client library to learn.
  3. Trace and evaluate
    Every request can be traced automatically. Review, evaluate, monitor, and iterate on AI applications and agents directly, no instrumentation or setup required.
  4. Scale as usage grows
    Pay per token as traffic increases, monitor your spend in the billing dashboard and set an inference budget. No capacity to forecast, no headroom to buy in advance.

Frequently asked questions

Which models are available through Serverless Inference?

Serverless Inference hosts a curated catalog of leading open-source models, updated as new models are released. If you need a specific model that isn't in the catalog, talk to us about your architecture.

Can I run my own fine-tuned model on Serverless Inference?

Yes. Bring your own LoRA weights and deploy them alongside the base catalog, without standing up or scaling dedicated serving infrastructure for every iteration. That matters most for teams moving quickly between training and inference: fine-tune, deploy, evaluate, and fine-tune again, all without a separate deployment step in between.

How is Serverless Inference billed?

Serverless Inference is billed per token as usage scales. There's no GPU-hour billing and no infrastructure to provision ahead of traffic, so cost tracks what you actually send and receive rather than capacity you reserved just in case.

Are there usage limits or concurrency caps I should plan around?

Yes. Serverless Inference applies a default spending cap by account tier and a concurrency limit per project to protect service quality and prevent unexpected charges. For current caps, concurrency limits, and how to request a change, see [Usage information and limits].

How does tracing and observability work?

We give enterprises full flexibility of data and observability when they operate on Serverless Inference. By default, Serverless Inference follows a Zero Data Retention policy. Enable tracing, and every request routed through Serverless Inference can be traced and observed via a simple built-in capability, so you can debug and monitor AI applications and agents without adding separate instrumentation.

Can I switch between models without redeploying anything?

Yes. Change the model parameter in your API call to switch models. Since every model in the catalog runs behind the same OpenAI-compatible endpoint, there's no redeployment, no new integration, and no new provider relationship to set up.

Can I evaluate a model before integrating it into my application?

Yes. Explore and compare models directly in the playground, with no endpoint to configure and no API key required to start. Once you've found the model that fits, point your application at the same endpoint you tested against.

Left
Right

Inference built on the Essential Cloud for AI

Serverless Inference gives you instant access to leading open-source models, per-token economics, and observability built in. Get started today.