Baseten vs SageMaker vs Vertex AI for Model Serving 2026
The short answer
Baseten is the faster path to production inference with real scale-to-zero; SageMaker and Vertex AI win when the model has to live inside your cloud account and your committed spend. Baseten bills per GPU-minute per replica and scales to zero on its free Basic plan. SageMaker can scale GPU endpoints to zero only through inference components, with several minutes of provisioning during which calls fail. Vertex AI — renamed Gemini Enterprise Agent Platform in 2026 — offers Scale To Zero only in preview. Facts below are from vendor documentation read October 8, 2026.
The comparison
| Baseten | Amazon SageMaker AI | Vertex AI / Gemini Enterprise Agent Platform | |
|---|---|---|---|
| What it is | Specialist inference platform | Full ML platform on AWS | Full ML and agent platform on Google Cloud |
| Billing | Per GPU-minute per replica, incl. deploy/scale-up minutes | Instance-hours (real-time); per-ms compute (serverless) | Node-hours for dedicated endpoints |
| Example GPU rate | H100 $0.10833/min (~$6.50/h); L4 $0.01414/min | Per instance type and region | Per machine type and region |
| Scale to zero (GPU) | Yes, standard | Real-time: only with inference components, minutes to provision, errors meanwhile; Async: yes | Scale To Zero in preview; 429 “model not yet ready” until a replica starts |
| Serverless GPU | — (dedicated per-minute replicas) | No — Serverless Inference has no GPUs, 6 GB max memory | No |
| Packaging | Truss (baseten model push) | Containers, JumpStart, inference components | Custom containers, Model Garden |
| Data residency / VPC | Enterprise: self-host in your cloud, custom regions | Native VPC, IAM, PrivateLink | Native VPC, IAM, VPC-SC |
| Compliance | SOC 2 Type II and HIPAA on Basic | Full AWS compliance catalogue | Full Google Cloud catalogue |
Platform notes
Baseten. Every deployment runs on an instance you pick by GPU and size (H100, L4:4x16, A10Gx4x16 and so on), defined in Truss config.yaml. You pay per minute per replica while it runs — including minutes spent deploying and scaling — and nothing when it has scaled to zero. The Basic plan is $0 a month, pay as you go, with dedicated deployments, Model APIs, training and fast cold starts; Pro adds priority access to scarce GPUs; Enterprise adds self-hosted deployments, custom regions and the ability to spend existing cloud commitments.
Amazon SageMaker AI. Three serving modes matter. Real-time endpoints run on instances you pay for by the hour; they can scale in to zero only if they host inference components and the variant’s MinInstanceCount is 0, and AWS warns that provisioning back from zero “takes several minutes” during which invocations return errors. Serverless Inference scales to zero and bills per millisecond, but does not support GPUs, tops out at 6 GB of memory and 200 concurrent invocations per endpoint — fine for small CPU models, not LLMs. Asynchronous inference queues requests and can scale to zero, which suits batch-like document or media jobs.
Vertex AI → Gemini Enterprise Agent Platform. Google renamed Vertex AI products in 2026: Vertex AI Endpoints are now “Gemini Enterprise Agent Platform Endpoints” and Online Inference keeps its function under the new name. Dedicated endpoints autoscale on CPU or GPU utilisation (default target 60%), with newer metrics such as vLLM KV-cache usage and queue depth in preview. By default minReplicaCount must be at least 1. Scale To Zero (preview) allows 0, scales down after an idle_scaledown_period such as 300 seconds, works only with one model per endpoint, not on shared public endpoints or multi-host GPU deployments, and answers 429 until a replica is ready — so clients need retry logic.
How to choose
- Startup or product team shipping a custom or fine-tuned model: Baseten. Least infrastructure work, real scale to zero.
- Regulated enterprise, data must stay in your AWS account: SageMaker real-time endpoints with inference components; keep one warm instance if latency matters.
- Already on Google Cloud with Gemini, agents and BigQuery: Gemini Enterprise Agent Platform dedicated endpoints; treat Scale To Zero as preview.
- Batch or queued workloads on AWS: SageMaker asynchronous inference.
Lock-in is mostly in the glue, not the model. All three serve standard containers running vLLM, SGLang or TensorRT-LLM, so the weights move; IAM, networking, monitoring and deployment pipelines do not. If you have no hyperscaler constraint, compare Baseten against other specialist platforms in best inference API vendors for autoscaling and fast cold starts, and see the broader Baseten vs Together AI vs Fireworks comparison for hosted open models.
Last verified: October 8, 2026. Rates are list USD; hyperscaler GPU prices vary by region and discount program.