Best Inference Providers for LoRA and Fine-Tuned Models 2026
The short answer
Per-token serverless hosting for LoRA adapters is mostly gone in 2026, so the choice is now which GPU platform serves your adapters best. Fireworks for managed multi-LoRA with the least work, Baseten for any custom model with per-minute billing and scale to zero, Together AI if you train there, Modal if traffic is spiky and you are happy running vLLM, and your own vLLM cluster at steady high volume. Prices are list USD as of October 7, 2026.
The comparison
| Provider | How LoRA is served | Adapters per deployment | GPU price (list) | Billing | Best for |
|---|---|---|---|---|---|
| Fireworks AI | Dedicated only: live merge (1 adapter, no overhead) or multi-LoRA addons | Many per base model; FP8/FP4 shapes don’t support addons | H100/H200 $8/hr · B200 $13/hr · B300 $15/hr | Per minute ($0.134/min H100) | Managed multi-LoRA, least ops |
| Together AI | Dedicated endpoint; attach adapters to a LoRA-enabled endpoint (beta) | Up to 16 for most models (5 for small Qwen3.5) | H100 $5.49/hr · B200 $8.99/hr on demand | Per GPU-hour | Teams fine-tuning on Together |
| Baseten | Your model in a Truss container; TensorRT-LLM LoRA swapping | Engine-dependent | H100 $0.10833/min ( | Per minute, scale to zero | Any custom model, non-LLMs too |
| Modal | You run vLLM/SGLang with LoRA enabled | Engine-dependent | H100 $0.001097/s ( | Per second, scale to zero | Spiky traffic, Python teams |
| Self-hosted vLLM | enable_lora=True, adapter chosen per request | Configurable | Your GPU cost | — | Steady high volume, full control |
Provider notes
Fireworks AI. Its docs (updated September 2026) give two paths. Live merge folds the adapter into the base weights at deploy time: one adapter per deployment, identical latency and throughput to the base model — the recommended choice for a single fine-tune. Multi-LoRA deploys the base model with addons enabled and loads adapters per request: some per-request overhead and lower peak throughput under high concurrency, but one GPU bill for many variants. You can import adapters trained elsewhere.
Together AI. Serverless LoRA and Serverless Multi-LoRA were discontinued; adapters now run on dedicated endpoints, and several adapters with the same base model can be attached to one LoRA-enabled endpoint and selected by model name. The endpoint must be private, LoRA-enabled and run the same base model. Good if you already fine-tune with Together, since LoRA is its default training method.
Baseten. Not LoRA-specific: you package any model (LLM, embedding, speech, diffusion) and pay per GPU-minute while replicas run, including while scaling. Its engineering blog describes serving many LoRA fine-tunes from one deployment using TensorRT-LLM’s adapter swapping. Choose it when the fine-tuned model is one of several custom models you need in production.
Modal. Serverless GPUs billed per second with scale to zero, $30 of free credits a month on the Starter plan, and you bring the serving engine. Two caveats from Modal’s pricing page: region selection costs 1.15–1.75x base prices and non-preemptible execution costs 3x, so the headline H100 rate is the floor.
Self-hosted vLLM. vLLM serves LoRA adapters per request on top of one base model with minimal overhead. It is the cheapest per token once a GPU is busy most of the day — and the most work.
How to choose
- One fine-tune, steady traffic: Fireworks live merge, or one Baseten deployment.
- Dozens of per-customer adapters on one base model: Fireworks multi-LoRA or a Together LoRA-enabled endpoint (check the 16-adapter cap).
- Traffic near zero at night: Baseten or Modal — both scale to zero.
- More than ~70% GPU utilisation around the clock: reserved GPUs and vLLM.
The cost trap is idle GPUs: a single H100 at $8 an hour left running is about $5,800 a month. Before fine-tuning, check whether a frontier model with a good prompt is cheaper — see current API prices.
Related
The full provider ranking is in best LLM inference providers; the three-way platform comparison is Baseten vs Together AI vs Fireworks; training itself is covered in how to fine-tune LLMs.
Last verified: October 7, 2026. GPU prices are list USD from vendor pricing pages.