AI agents · OpenClaw · self-hosting · automation

Quick Answer

Baseten vs SageMaker vs Vertex AI for Model Serving 2026

Published:

The short answer

Baseten is the faster path to production inference with real scale-to-zero; SageMaker and Vertex AI win when the model has to live inside your cloud account and your committed spend. Baseten bills per GPU-minute per replica and scales to zero on its free Basic plan. SageMaker can scale GPU endpoints to zero only through inference components, with several minutes of provisioning during which calls fail. Vertex AI — renamed Gemini Enterprise Agent Platform in 2026 — offers Scale To Zero only in preview. Facts below are from vendor documentation read October 8, 2026.

The comparison

BasetenAmazon SageMaker AIVertex AI / Gemini Enterprise Agent Platform
What it isSpecialist inference platformFull ML platform on AWSFull ML and agent platform on Google Cloud
BillingPer GPU-minute per replica, incl. deploy/scale-up minutesInstance-hours (real-time); per-ms compute (serverless)Node-hours for dedicated endpoints
Example GPU rateH100 $0.10833/min (~$6.50/h); L4 $0.01414/minPer instance type and regionPer machine type and region
Scale to zero (GPU)Yes, standardReal-time: only with inference components, minutes to provision, errors meanwhile; Async: yesScale To Zero in preview; 429 “model not yet ready” until a replica starts
Serverless GPU— (dedicated per-minute replicas)No — Serverless Inference has no GPUs, 6 GB max memoryNo
PackagingTruss (baseten model push)Containers, JumpStart, inference componentsCustom containers, Model Garden
Data residency / VPCEnterprise: self-host in your cloud, custom regionsNative VPC, IAM, PrivateLinkNative VPC, IAM, VPC-SC
ComplianceSOC 2 Type II and HIPAA on BasicFull AWS compliance catalogueFull Google Cloud catalogue

Platform notes

Baseten. Every deployment runs on an instance you pick by GPU and size (H100, L4:4x16, A10Gx4x16 and so on), defined in Truss config.yaml. You pay per minute per replica while it runs — including minutes spent deploying and scaling — and nothing when it has scaled to zero. The Basic plan is $0 a month, pay as you go, with dedicated deployments, Model APIs, training and fast cold starts; Pro adds priority access to scarce GPUs; Enterprise adds self-hosted deployments, custom regions and the ability to spend existing cloud commitments.

Amazon SageMaker AI. Three serving modes matter. Real-time endpoints run on instances you pay for by the hour; they can scale in to zero only if they host inference components and the variant’s MinInstanceCount is 0, and AWS warns that provisioning back from zero “takes several minutes” during which invocations return errors. Serverless Inference scales to zero and bills per millisecond, but does not support GPUs, tops out at 6 GB of memory and 200 concurrent invocations per endpoint — fine for small CPU models, not LLMs. Asynchronous inference queues requests and can scale to zero, which suits batch-like document or media jobs.

Vertex AI → Gemini Enterprise Agent Platform. Google renamed Vertex AI products in 2026: Vertex AI Endpoints are now “Gemini Enterprise Agent Platform Endpoints” and Online Inference keeps its function under the new name. Dedicated endpoints autoscale on CPU or GPU utilisation (default target 60%), with newer metrics such as vLLM KV-cache usage and queue depth in preview. By default minReplicaCount must be at least 1. Scale To Zero (preview) allows 0, scales down after an idle_scaledown_period such as 300 seconds, works only with one model per endpoint, not on shared public endpoints or multi-host GPU deployments, and answers 429 until a replica is ready — so clients need retry logic.

How to choose

  1. Startup or product team shipping a custom or fine-tuned model: Baseten. Least infrastructure work, real scale to zero.
  2. Regulated enterprise, data must stay in your AWS account: SageMaker real-time endpoints with inference components; keep one warm instance if latency matters.
  3. Already on Google Cloud with Gemini, agents and BigQuery: Gemini Enterprise Agent Platform dedicated endpoints; treat Scale To Zero as preview.
  4. Batch or queued workloads on AWS: SageMaker asynchronous inference.

Lock-in is mostly in the glue, not the model. All three serve standard containers running vLLM, SGLang or TensorRT-LLM, so the weights move; IAM, networking, monitoring and deployment pipelines do not. If you have no hyperscaler constraint, compare Baseten against other specialist platforms in best inference API vendors for autoscaling and fast cold starts, and see the broader Baseten vs Together AI vs Fireworks comparison for hosted open models.

Last verified: October 8, 2026. Rates are list USD; hyperscaler GPU prices vary by region and discount program.

Sources