AI agents · OpenClaw · self-hosting · automation

Quick Answer

Which Local AI Models Can Run on a MacBook Air? (2026)

Published:

The short answer

A MacBook Air runs small and mixture-of-experts models well, and dense 30B-class models only slowly. On the M5 Air (16GB standard, 24GB or 32GB optional, 153GB/s memory bandwidth): Gemma 4 E4B on 16GB, Gemma 4 26B A4B on 24GB or more, and Qwen3.8-27B or Gemma 4 31B on 32GB if you can accept single-digit tokens per second. Facts as of October 7, 2026.

What fits, by memory

Air memoryBest models (4-bit)Approx. weights in RAMFeels like
16GBGemma 4 E4B · Gemma 4 E2B · other 7–8B models~3–6GBFast; good for chat, summaries, private notes
24GBGemma 4 26B A4B (MoE) · everything above~15–17GBFast for its quality — only ~4B parameters active per token
32GBQwen3.8-27B · Gemma 4 31B (dense) · everything above~16–19GBStrong answers, slow output; fine for batch jobs

Leave 6–8GB for macOS and your apps. Weight sizes are our estimates for 4-bit quantised files (about 0.55–0.6GB per billion parameters); the exact download depends on the quantisation you choose.

Why speed falls off on dense models

Generating each token means reading the active weights from memory once, so a rough ceiling is bandwidth ÷ active weight size. At 153GB/s, a dense 27B model at ~16GB tops out around 9 tokens per second before any overhead; Gemma 4 26B A4B reads only ~4B active parameters per token, so it runs several times faster despite a similar file size. That is why the 26B MoE is the sweet spot on an Air. The Air is also fanless: long generations make it throttle sooner than a MacBook Pro.

The models

  • Gemma 4 E2B / E4B (Google, April 2, 2026; Apache 2.0). Built for on-device use: 128K context, image and video input, and native audio input for speech. The default for a 16GB Air.
  • Gemma 4 26B A4B (MoE) and Gemma 4 31B (dense). Up to 256K context; Google says the 31B ranked #3 and the 26B #6 among open models on the Arena text leaderboard at launch.
  • Qwen3.8-27B (Alibaba, open weights since mid-August 2026). A dense vision-language model with thinking control; the strongest coding and agent model that fits a 32GB Air, at the cost of speed.

Larger open models such as Kimi K3 or DeepSeek V4.1 Flash need far more memory than any MacBook Air has; use them through an API (see current API prices).

How to run them

  1. Install LM Studio (graphical, supports Apple’s MLX format) or Ollama (command line plus a local API on port 11434).
  2. Download the model at 4-bit (Q4_K_M in GGUF, or 4-bit MLX).
  3. Start with a 8K–32K context window; long contexts consume memory on top of the weights.
  4. Close heavy apps before loading a 20GB+ model, and keep the Air plugged in for long runs.

Should you buy a MacBook Air for local AI?

Buy 32GB if local models matter at all — it is the only Air configuration that loads the 27B–31B class. Apple’s US store listed MacBook Air M5 configurations from $1,299 on October 7, 2026 (it launched at $1,099 in March and was raised in June). If local models are your main workload, a MacBook Pro with higher bandwidth and active cooling is the better tool.

Best local LLMs for Mac · best small LLMs for a single GPU · what hardware you need to run LLMs locally · Ollama MLX vs Metal on Apple silicon.

Last verified: October 7, 2026.

Sources