How to Self-Host an LLM for Production Use (2026 Guide)
The short answer
Self-hosting an LLM for production is seven steps: choose the model, size the GPUs, pick a serving engine, deploy behind a gateway, autoscale, monitor, and gate changes on evals. The default stack in October 2026 is an open-weight model served by vLLM (v0.31.0, released October 5, 2026) or SGLang on rented or owned NVIDIA GPUs, behind an OpenAI-compatible endpoint so application code doesn’t care whether the model is yours or a vendor’s.
Step 1 — Choose the model and check the licence
Start from the smallest model that passes your eval set, not the biggest you can fit. Check the licence for commercial use (MIT and Apache-2.0 are simplest; some open models carry custom terms) and whether weights are actually released — Mistral Large 4, for example, is in API preview with weights promised for the end of October 2026. Current options are in the best open-weight models to self-host.
Done when: you have one primary model, one smaller fallback, and a written licence check.
Step 2 — Size the GPUs
Weights ≈ parameters × bytes per parameter: about 2 bytes in BF16, 1 in FP8, 0.5 at 4-bit. Add 20–50% for KV cache, more for long contexts and many concurrent users.
| Model class | Precision | Weights | Typical hardware | RunPod on-demand (Oct 2026) |
|---|---|---|---|---|
| 7–14B | FP8 | 7–14 GB | L4 24 GB or L40S 48 GB | L4 $0.49/h · L40S $1.09/h |
| 30–40B | FP8 | 30–40 GB | L40S 48 GB or H100 80 GB | H100 SXM $3.49/h |
| 100–120B | 4-bit | 50–60 GB | H100 80 GB or RTX Pro 6000 96 GB | RTX Pro 6000 $2.09/h |
| 300B+ MoE | FP8 / FP4 | 300 GB+ | 8× H200 or B200 node | H200 $4.59/h · B200 $6.79/h per GPU |
Done when: the model plus KV cache for your target context and concurrency fits with 10–15% headroom.
Step 3 — Pick the serving engine
Use vLLM as the default: continuous batching, paged attention, prefix caching, tensor and expert parallelism, quantized formats and an OpenAI-compatible server. Version 0.31.0 (717 commits from 307 contributors) added DeepSeek-V4.1-Flash kernels and a vllm preload weight-cache daemon that keeps weights resident in GPU memory across engine restarts. SGLang is the main alternative, strong on structured output and prefix-heavy agent workloads, and also ships day-one support for new open models. Ollama and LM Studio are for development, not multi-user serving — see vLLM vs Ollama vs LM Studio.
A minimal production launch looks like:
vllm serve <org>/<model> \
--tensor-parallel-size 2 \
--max-model-len 65536 \
--gpu-memory-utilization 0.90 \
--api-key "$VLLM_API_KEY"
Done when: the server answers an OpenAI-format /v1/chat/completions request from your app.
Step 4 — Put a gateway in front
Never expose the engine directly. Terminate TLS, authenticate per client or tenant, enforce rate and token limits, log requests without storing sensitive prompts, and route to a hosted API when your cluster is saturated or down. Keeping the API OpenAI-compatible means failover is a base-URL change — see how to add LLM failover.
Done when: a request without a valid key is rejected and a forced outage fails over cleanly.
Step 5 — Autoscale on the right signal
GPU servers saturate on KV-cache memory and queue depth long before CPU. Scale on requests waiting or cache usage; Google’s managed endpoints, for example, now expose vLLM KV-cache usage and waiting-request count as autoscaling metrics. Keep at least one warm replica during business hours — loading weights for a large model takes minutes. If traffic is spiky, a scale-to-zero GPU platform may beat your own cluster: see best inference vendors for autoscaling and cold starts.
Done when: a load test at 2× expected peak keeps time-to-first-token inside your SLO.
Step 6 — Monitor
Scrape the engine’s Prometheus /metrics endpoint and alert on time to first token, inter-token latency, tokens per second, queue length, KV-cache utilisation, error rate and GPU memory. Track cost per million tokens as utilisation × GPU price ÷ tokens served.
Done when: you get paged for a latency regression before users report it.
Step 7 — Gate every change on evals
Engine upgrades, quantization changes and new checkpoints all change outputs. Keep a fixed eval set from real traffic, run it on every change, and roll out behind a percentage canary with the previous version kept warm.
Done when: no model, engine or config change reaches 100% of traffic without passing the eval set.
Check the economics before you start
One H100 at $3.49 an hour is about $2,550 a month at any utilisation; DeepSeek V4.1 Flash costs $0.15/$0.60 per million tokens off-peak through DeepSeek’s API. Self-hosting usually wins on data control, latency and customisation, rarely on price — run the self-host vs API break-even with your real volume and compare current API prices.
Last verified: October 8, 2026. GPU prices are RunPod list rates on that date.