AI agents · OpenClaw · self-hosting · automation

Quick Answer

Best Inference Providers for LoRA and Fine-Tuned Models 2026

Published:

The short answer

Per-token serverless hosting for LoRA adapters is mostly gone in 2026, so the choice is now which GPU platform serves your adapters best. Fireworks for managed multi-LoRA with the least work, Baseten for any custom model with per-minute billing and scale to zero, Together AI if you train there, Modal if traffic is spiky and you are happy running vLLM, and your own vLLM cluster at steady high volume. Prices are list USD as of October 7, 2026.

The comparison

ProviderHow LoRA is servedAdapters per deploymentGPU price (list)BillingBest for
Fireworks AIDedicated only: live merge (1 adapter, no overhead) or multi-LoRA addonsMany per base model; FP8/FP4 shapes don’t support addonsH100/H200 $8/hr · B200 $13/hr · B300 $15/hrPer minute ($0.134/min H100)Managed multi-LoRA, least ops
Together AIDedicated endpoint; attach adapters to a LoRA-enabled endpoint (beta)Up to 16 for most models (5 for small Qwen3.5)H100 $5.49/hr · B200 $8.99/hr on demandPer GPU-hourTeams fine-tuning on Together
BasetenYour model in a Truss container; TensorRT-LLM LoRA swappingEngine-dependentH100 $0.10833/min ($6.50/hr) · B200 $0.16633/min ($9.98/hr)Per minute, scale to zeroAny custom model, non-LLMs too
ModalYou run vLLM/SGLang with LoRA enabledEngine-dependentH100 $0.001097/s ($3.95/hr) · B200 $0.001736/s ($6.25/hr)Per second, scale to zeroSpiky traffic, Python teams
Self-hosted vLLMenable_lora=True, adapter chosen per requestConfigurableYour GPU cost—Steady high volume, full control

Provider notes

Fireworks AI. Its docs (updated September 2026) give two paths. Live merge folds the adapter into the base weights at deploy time: one adapter per deployment, identical latency and throughput to the base model — the recommended choice for a single fine-tune. Multi-LoRA deploys the base model with addons enabled and loads adapters per request: some per-request overhead and lower peak throughput under high concurrency, but one GPU bill for many variants. You can import adapters trained elsewhere.

Together AI. Serverless LoRA and Serverless Multi-LoRA were discontinued; adapters now run on dedicated endpoints, and several adapters with the same base model can be attached to one LoRA-enabled endpoint and selected by model name. The endpoint must be private, LoRA-enabled and run the same base model. Good if you already fine-tune with Together, since LoRA is its default training method.

Baseten. Not LoRA-specific: you package any model (LLM, embedding, speech, diffusion) and pay per GPU-minute while replicas run, including while scaling. Its engineering blog describes serving many LoRA fine-tunes from one deployment using TensorRT-LLM’s adapter swapping. Choose it when the fine-tuned model is one of several custom models you need in production.

Modal. Serverless GPUs billed per second with scale to zero, $30 of free credits a month on the Starter plan, and you bring the serving engine. Two caveats from Modal’s pricing page: region selection costs 1.15–1.75x base prices and non-preemptible execution costs 3x, so the headline H100 rate is the floor.

Self-hosted vLLM. vLLM serves LoRA adapters per request on top of one base model with minimal overhead. It is the cheapest per token once a GPU is busy most of the day — and the most work.

How to choose

  1. One fine-tune, steady traffic: Fireworks live merge, or one Baseten deployment.
  2. Dozens of per-customer adapters on one base model: Fireworks multi-LoRA or a Together LoRA-enabled endpoint (check the 16-adapter cap).
  3. Traffic near zero at night: Baseten or Modal — both scale to zero.
  4. More than ~70% GPU utilisation around the clock: reserved GPUs and vLLM.

The cost trap is idle GPUs: a single H100 at $8 an hour left running is about $5,800 a month. Before fine-tuning, check whether a frontier model with a good prompt is cheaper — see current API prices.

The full provider ranking is in best LLM inference providers; the three-way platform comparison is Baseten vs Together AI vs Fireworks; training itself is covered in how to fine-tune LLMs.

Last verified: October 7, 2026. GPU prices are list USD from vendor pricing pages.

Sources