What Hardware Do I Need to Run LLMs Locally? (2026)
The short answer
Two numbers decide what you can run: how much memory the model fits in, and how fast that memory is. Capacity decides whether a model loads; bandwidth decides how many tokens per second you get. Buy for the biggest model you actually need, then for bandwidth. Compute (TFLOPS) matters mostly for prompt processing and fine-tuning.
Step 1: Work out the memory you need
Rule of thumb for the weights: parameters (in billions) × bits per weight ÷ 8 = gigabytes. Add 10–30% for the KV cache (context) and runtime overhead; long contexts need more.
| Model size | 4-bit (Q4) weights | Plan for | 8-bit weights |
|---|---|---|---|
| 3–4B | ~2 GB | 4–8 GB | ~4 GB |
| 8B | ~4–5 GB | 8 GB | ~8 GB |
| 27–32B | ~14–17 GB | 20–24 GB | ~30 GB |
| 70B | ~35–40 GB | 48 GB | ~70 GB |
| 120B | ~60–65 GB | 80 GB | ~120 GB |
| 300B+ (large MoE) | ~150 GB+ | 192–256 GB | ~300 GB+ |
Mixture-of-experts models need memory for all their parameters, even though only a fraction are active per token. Xiaomi’s MiMo-V2.6-Flash (310B total, 15B active) needs roughly 155 GB at 4-bit; DeepSeek V4.1 Flash (552B total) roughly 280 GB. Their small active count is what makes them fast once loaded.
Step 2: Work out the speed you’ll get
Generating each token reads the active weights from memory once, so a rough ceiling is memory bandwidth ÷ size of the active weights. A 27B dense model at 4-bit (~16 GB) on 273 GB/s tops out near 17 tokens per second; on 614 GB/s near 38; on 1.2 TB/s near 75. Real numbers land below these ceilings. MoE models with few active parameters run much faster than their total size suggests.
Step 3: Match it to hardware (October 2026 prices)
| Hardware | Memory | Bandwidth | Price | Runs comfortably |
|---|---|---|---|---|
| Any 16 GB Apple Silicon Mac or 12–16 GB GPU | 16 GB | varies | what you own | 3B–14B at 4-bit |
| Mac mini M5 Pro | 24 GB base, up to 64 GB | — | $1,699 (24 GB); $2,699 (64 GB) | up to ~70B at 4-bit (64 GB) |
| NVIDIA DGX Spark 64GB (partner systems from Oct 23, 2026) | 64 GB unified | 273 GB/s | from $4,999 | models up to ~100B (NVIDIA); CUDA |
| NVIDIA DGX Spark 128GB | 128 GB unified | 273 GB/s | $6,950 | up to ~200B; two clustered 64 GB units also reach 128 GB |
| Mac Studio M5 Max | up to 128 GB | up to 614 GB/s | from $2,499 | 70B–120B at 4-bit, faster than Spark |
| Mac Studio M5 Ultra | up to 512 GB (512 GB config ships late October 2026) | 1.2 TB/s | from $5,499 | 300B+ MoE models on one machine |
| Workstation NVIDIA GPU (24–96 GB VRAM) | VRAM | up to ~1.8 TB/s | varies | fastest on models that fit in VRAM |
Notes:
- DGX Spark’s price moved in October 2026. NVIDIA announced the 64GB configuration on October 2, 2026 at $4,999 from Acer, ASUS, Dell, Gigabyte, HP and MSI, while the 128GB model now sells for $6,950 (it launched at $3,999 in October 2025). Buy it for CUDA compatibility, not speed: its 273 GB/s is the slowest memory here.
- Apple’s numbers: Mac Studio with M5 Max has an 18-core CPU, up to a 40-core GPU and up to 128 GB; M5 Ultra scales to a 36-core CPU, 80-core GPU, 512 GB and 1.2 TB/s. The Mac mini M5 Pro configures up to 64 GB.
- Discrete GPUs win on tokens per second but cap model size at their VRAM; splitting a model across two cards works in llama.cpp and vLLM but adds complexity.
Step 4: Pick the software
Ollama and LM Studio are the easiest way to start on Mac, Windows and Linux; both use Metal on Apple Silicon and CUDA on NVIDIA. Use vLLM when you serve several users from an NVIDIA box. Use MLX-based tools on a Mac for the best Apple Silicon performance.
The pick
- Trying local models: use the computer you have with an 8B model in Ollama.
- Daily coding or private assistant on 27–70B models: Mac mini M5 Pro with 64 GB, or Mac Studio M5 Max.
- CUDA-only stack or agent development: DGX Spark 64GB.
- The largest open models on one desk: Mac Studio M5 Ultra with 256–512 GB.
Before spending $5,000, compare with API prices: frontier open-weight models cost cents per task through inference providers. Local hardware pays off for privacy, sustained heavy use or offline work — not usually on cost alone.
Related: best hardware for local AI agents, vLLM vs Ollama vs LM Studio, cheapest inference provider for open-weight models.
Last verified: October 6, 2026. Speed figures are bandwidth-based ceilings, not benchmarks; prices are U.S. list prices.