AI agents · OpenClaw · self-hosting · automation

Quick Answer

Cheapest AI API August 2026: Price Per Token Ranked

Published:

The Short Answer

The cheapest AI API in August 2026 is DeepSeek V4-Flash ($0.14/$0.28) at the frontier-adjacent tier. For lighter tasks, Amazon Nova Micro, Qwen 3.7 Flash, and Gemini 3.5 Flash-Lite go even lower.

The Ranking (per million tokens)

RankModelInputOutputTier
1Llama 3.1 8B Instruct$0.02~$0.03Small/cheapest overall
2Qwen 3.7 Flash$0.03$0.13Light coding
3Amazon Nova Micro$0.035$0.14Light/fast
4Amazon Nova Lite$0.06$0.24Light
5DeepSeek V4-Flash$0.14$0.28Frontier-adjacent
6Gemini 3.5 Flash-Lite$0.30$2.50Fast hosted
Gemini 3.6 Flash$1.50$7.50Value workhorse

How To Read This

  • Best value at the frontier → DeepSeek V4-Flash. The “price floor of the frontier-adjacent market” — roughly 10x-90x cheaper than Claude Opus 5 and GPT-5.6 Sol. Cache-hit input is just $0.0028/MTok, so cache-heavy agent loops cost almost nothing.
  • Absolute lowest for light tasks → Llama 3.1 8B, Qwen 3.7 Flash, and Amazon Nova Micro — great for classification, extraction, and short responses, but not frontier reasoning.
  • Cheap hosted convenience → Gemini 3.5 Flash-Lite ($0.30/$2.50) and Gemini 3.6 Flash ($1.50/$7.50), which uses up to 65% fewer output tokens.

Watch Outs

  • DeepSeek peak pricing: a 2x peak-hour surcharge (Beijing-time windows) has been announced but has no effective date as of early August 2026 — standard rates still apply. The deepseek-chat / deepseek-reasoner names were retired July 24, 2026.
  • Cheap ≠ capable: Nova Micro and Llama 8B are small models. For hard multi-step tasks, DeepSeek V4-Flash or a flagship is worth the higher price.

Verdict

  • Cheapest frontier-adjacent → DeepSeek V4-Flash
  • Cheapest overall → Llama 3.1 8B / Qwen 3.7 Flash
  • Cheapest hosted mid-tier → Gemini 3.5 Flash-Lite

Sources