AI agents · OpenClaw · self-hosting · automation

Quick Answer

Best Inference API Vendors for Autoscaling & Cold Starts

Published:

The short answer

First check whether you need your own GPU at all. If the open model you want (DeepSeek V4.1 Flash, GLM-5.3, gpt-oss, Qwen 3.8) is in a provider’s catalog, a per-token serverless API from Fireworks, Together AI or DeepInfra has no cold start, autoscales for you and needs zero deployment — see the cheapest inference provider for open-weight models. You need a GPU platform only for custom, fine-tuned or non-catalog models. Then pick by how you work: Modal for Python-native code, RunPod Serverless for the lowest rates, Baseten for production autoscaling, Cerebrium for guaranteed burst capacity, Replicate for the simplest workflow. Prices are list USD from vendor pages, read October 8, 2026.

The comparison

VendorBillingH100 rate (list)Scale to zeroCold-start featuresFree tier / planSimplest deploy path
ModalPer second~$3.95/hYesWeight caching on volumes; you control the imageStarter $0 + $30/mo credit, 10 GPU concurrency; Team $250/mo + $100 credit, 50 GPUsPython decorators, modal deploy
RunPod ServerlessPer second$4.79/h flex worker (B200 $8.64, A100 $2.72, L4 $0.69)Yes (flex workers)FlashBoot, “as low as 500 ms” for busy endpoints, no extra costPay as you goDocker image or template endpoint
BasetenPer minute, including deploy and scale-up minutes$0.10833/min (~$6.50/h)Yes”Fast cold starts” on Basic; TensorRT-LLM enginesBasic $0/mo pay as you go; Pro/Enterprise by quoteTruss config + baseten model push
CerebriumPer second, by allocated GPU/CPU/RAMNot published on page; L4 example $0.000257/s with 2 vCPU, 10 GBYesContainers scale in 1–3 s; memory and GPU snapshottingHobby free (5 GPUs); Standard $100/mo (30 GPUs)Python app + CLI
ReplicatePer second$0.001525/s (~$5.49/h); L40S $0.000975/s; A100 80GB $0.0014/sFast-booting fine-tunes onlyPublic models are warm; private models bill setup + idle timePay as you goCog container, push and call

Vendor notes

Modal. Serverless GPUs billed per second with automatic scale up and down. Its pricing page argues the case directly: 75 reserved GPUs at $3 per hour cost $5,400 a day, while an average of 50 Modal GPUs at $3.95 per hour costs $4,740 — 12% less, with no degraded service at peak. Two multipliers to budget for: region selection costs 1.15–1.75x base prices and non-preemptible execution 3x, so the headline rate is the floor.

RunPod Serverless. The cheapest published per-second rates of the group, and the widest GPU menu (16 GB cards at $0.58 per hour up to B300 at $9.98). Flex workers scale to zero; active workers stay warm at a discount negotiated with sales. FlashBoot keeps workers that have just served traffic ready so popular endpoints restart in as little as 500 ms by RunPod’s own measurement — an endpoint called once a day will still see a full cold start.

Baseten. Built for production inference rather than experiments: per-minute billing per replica, scale to zero, autoscaling, and SOC 2 Type II and HIPAA on the free Basic plan. Note that you pay for the minutes a replica spends deploying or scaling, so aggressive scale-to-zero on a large model can cost more than one warm replica. Pro adds priority access to scarce GPUs; Enterprise can run in your own cloud and use existing cloud commitments.

Cerebrium. Bills the resources you allocate only while they run, and says containers scale in 1–3 seconds with memory and GPU snapshotting to restore faster. Its distinctive offer is guaranteed burst capacity without a 24/7 reservation — for example access to up to 50 H100s at any time against a $10,000 minimum monthly spend.

Replicate. The fastest path from a model to an HTTP endpoint. The catch for spiky traffic: most private models run on dedicated hardware and bill setup and idle time as well as active time; only “fast booting fine-tunes” bill active time alone.

How to choose

  1. Model is in a catalog: use a per-token API. No GPUs, no cold starts.
  2. Custom model, traffic in bursts with idle hours: RunPod Serverless or Modal.
  3. Customer-facing latency SLO: Baseten or Cerebrium with one warm minimum replica in business hours, scale to zero at night.
  4. Busy more than ~75% of the day: dedicated GPUs — RunPod’s own H100 SXM pod is $3.49 per hour against $4.79 serverless.

Avoid hyperscalers for this job. Amazon SageMaker Serverless Inference does not support GPUs (6 GB memory maximum), and real-time endpoints that scale to zero take several minutes to provision, returning errors meanwhile. Vertex AI — renamed Gemini Enterprise Agent Platform — offers Scale To Zero only in preview and answers HTTP 429 until a replica starts. The full trade-off is in Baseten vs SageMaker vs Vertex AI.

The full provider ranking is best LLM inference providers; serving fine-tunes is covered in best inference providers for LoRA models.

Last verified: October 8, 2026. GPU prices are list USD from vendor pricing pages and change often.

Sources