Which Local AI Models Can Run on a MacBook Air? (2026)
The short answer
A MacBook Air runs small and mixture-of-experts models well, and dense 30B-class models only slowly. On the M5 Air (16GB standard, 24GB or 32GB optional, 153GB/s memory bandwidth): Gemma 4 E4B on 16GB, Gemma 4 26B A4B on 24GB or more, and Qwen3.8-27B or Gemma 4 31B on 32GB if you can accept single-digit tokens per second. Facts as of October 7, 2026.
What fits, by memory
| Air memory | Best models (4-bit) | Approx. weights in RAM | Feels like |
|---|---|---|---|
| 16GB | Gemma 4 E4B · Gemma 4 E2B · other 7–8B models | ~3–6GB | Fast; good for chat, summaries, private notes |
| 24GB | Gemma 4 26B A4B (MoE) · everything above | ~15–17GB | Fast for its quality — only ~4B parameters active per token |
| 32GB | Qwen3.8-27B · Gemma 4 31B (dense) · everything above | ~16–19GB | Strong answers, slow output; fine for batch jobs |
Leave 6–8GB for macOS and your apps. Weight sizes are our estimates for 4-bit quantised files (about 0.55–0.6GB per billion parameters); the exact download depends on the quantisation you choose.
Why speed falls off on dense models
Generating each token means reading the active weights from memory once, so a rough ceiling is bandwidth ÷ active weight size. At 153GB/s, a dense 27B model at ~16GB tops out around 9 tokens per second before any overhead; Gemma 4 26B A4B reads only ~4B active parameters per token, so it runs several times faster despite a similar file size. That is why the 26B MoE is the sweet spot on an Air. The Air is also fanless: long generations make it throttle sooner than a MacBook Pro.
The models
- Gemma 4 E2B / E4B (Google, April 2, 2026; Apache 2.0). Built for on-device use: 128K context, image and video input, and native audio input for speech. The default for a 16GB Air.
- Gemma 4 26B A4B (MoE) and Gemma 4 31B (dense). Up to 256K context; Google says the 31B ranked #3 and the 26B #6 among open models on the Arena text leaderboard at launch.
- Qwen3.8-27B (Alibaba, open weights since mid-August 2026). A dense vision-language model with thinking control; the strongest coding and agent model that fits a 32GB Air, at the cost of speed.
Larger open models such as Kimi K3 or DeepSeek V4.1 Flash need far more memory than any MacBook Air has; use them through an API (see current API prices).
How to run them
- Install LM Studio (graphical, supports Apple’s MLX format) or Ollama (command line plus a local API on port 11434).
- Download the model at 4-bit (Q4_K_M in GGUF, or 4-bit MLX).
- Start with a 8K–32K context window; long contexts consume memory on top of the weights.
- Close heavy apps before loading a 20GB+ model, and keep the Air plugged in for long runs.
Should you buy a MacBook Air for local AI?
Buy 32GB if local models matter at all — it is the only Air configuration that loads the 27B–31B class. Apple’s US store listed MacBook Air M5 configurations from $1,299 on October 7, 2026 (it launched at $1,099 in March and was raised in June). If local models are your main workload, a MacBook Pro with higher bandwidth and active cooling is the better tool.
Related
Best local LLMs for Mac · best small LLMs for a single GPU · what hardware you need to run LLMs locally · Ollama MLX vs Metal on Apple silicon.
Last verified: October 7, 2026.