AI agents · OpenClaw · self-hosting · automation

Quick Answer

What Hardware Do I Need to Run LLMs Locally? (2026)

Published:

The short answer

Two numbers decide what you can run: how much memory the model fits in, and how fast that memory is. Capacity decides whether a model loads; bandwidth decides how many tokens per second you get. Buy for the biggest model you actually need, then for bandwidth. Compute (TFLOPS) matters mostly for prompt processing and fine-tuning.

Step 1: Work out the memory you need

Rule of thumb for the weights: parameters (in billions) × bits per weight ÷ 8 = gigabytes. Add 10–30% for the KV cache (context) and runtime overhead; long contexts need more.

Model size4-bit (Q4) weightsPlan for8-bit weights
3–4B~2 GB4–8 GB~4 GB
8B~4–5 GB8 GB~8 GB
27–32B~14–17 GB20–24 GB~30 GB
70B~35–40 GB48 GB~70 GB
120B~60–65 GB80 GB~120 GB
300B+ (large MoE)~150 GB+192–256 GB~300 GB+

Mixture-of-experts models need memory for all their parameters, even though only a fraction are active per token. Xiaomi’s MiMo-V2.6-Flash (310B total, 15B active) needs roughly 155 GB at 4-bit; DeepSeek V4.1 Flash (552B total) roughly 280 GB. Their small active count is what makes them fast once loaded.

Step 2: Work out the speed you’ll get

Generating each token reads the active weights from memory once, so a rough ceiling is memory bandwidth ÷ size of the active weights. A 27B dense model at 4-bit (~16 GB) on 273 GB/s tops out near 17 tokens per second; on 614 GB/s near 38; on 1.2 TB/s near 75. Real numbers land below these ceilings. MoE models with few active parameters run much faster than their total size suggests.

Step 3: Match it to hardware (October 2026 prices)

HardwareMemoryBandwidthPriceRuns comfortably
Any 16 GB Apple Silicon Mac or 12–16 GB GPU16 GBvarieswhat you own3B–14B at 4-bit
Mac mini M5 Pro24 GB base, up to 64 GB—$1,699 (24 GB); $2,699 (64 GB)up to ~70B at 4-bit (64 GB)
NVIDIA DGX Spark 64GB (partner systems from Oct 23, 2026)64 GB unified273 GB/sfrom $4,999models up to ~100B (NVIDIA); CUDA
NVIDIA DGX Spark 128GB128 GB unified273 GB/s$6,950up to ~200B; two clustered 64 GB units also reach 128 GB
Mac Studio M5 Maxup to 128 GBup to 614 GB/sfrom $2,49970B–120B at 4-bit, faster than Spark
Mac Studio M5 Ultraup to 512 GB (512 GB config ships late October 2026)1.2 TB/sfrom $5,499300B+ MoE models on one machine
Workstation NVIDIA GPU (24–96 GB VRAM)VRAMup to ~1.8 TB/svariesfastest on models that fit in VRAM

Notes:

  • DGX Spark’s price moved in October 2026. NVIDIA announced the 64GB configuration on October 2, 2026 at $4,999 from Acer, ASUS, Dell, Gigabyte, HP and MSI, while the 128GB model now sells for $6,950 (it launched at $3,999 in October 2025). Buy it for CUDA compatibility, not speed: its 273 GB/s is the slowest memory here.
  • Apple’s numbers: Mac Studio with M5 Max has an 18-core CPU, up to a 40-core GPU and up to 128 GB; M5 Ultra scales to a 36-core CPU, 80-core GPU, 512 GB and 1.2 TB/s. The Mac mini M5 Pro configures up to 64 GB.
  • Discrete GPUs win on tokens per second but cap model size at their VRAM; splitting a model across two cards works in llama.cpp and vLLM but adds complexity.

Step 4: Pick the software

Ollama and LM Studio are the easiest way to start on Mac, Windows and Linux; both use Metal on Apple Silicon and CUDA on NVIDIA. Use vLLM when you serve several users from an NVIDIA box. Use MLX-based tools on a Mac for the best Apple Silicon performance.

The pick

  • Trying local models: use the computer you have with an 8B model in Ollama.
  • Daily coding or private assistant on 27–70B models: Mac mini M5 Pro with 64 GB, or Mac Studio M5 Max.
  • CUDA-only stack or agent development: DGX Spark 64GB.
  • The largest open models on one desk: Mac Studio M5 Ultra with 256–512 GB.

Before spending $5,000, compare with API prices: frontier open-weight models cost cents per task through inference providers. Local hardware pays off for privacy, sustained heavy use or offline work — not usually on cost alone.

Related: best hardware for local AI agents, vLLM vs Ollama vs LM Studio, cheapest inference provider for open-weight models.

Last verified: October 6, 2026. Speed figures are bandwidth-based ceilings, not benchmarks; prices are U.S. list prices.

Sources