Quick Answer
Cheapest AI API 2026: Price-Per-Token Ranked
The Short Answer
The cheapest capable AI API in 2026 is DeepSeek V4 Flash at $0.14/$0.28 per million tokens (cache-hit input $0.0028). For a bit more ceiling, Gemini 3.6 Flash ($1.50/$7.50), GPT-5.6 Luna ($1/$6), and Grok 4.5 ($2/$6) are the value frontier picks.
Ranked by Price (input/output per MTok)
| Model | Input | Output | Notes |
|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.28 | Cheapest; cache-hit in $0.0028; 2x peak pending |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Jul 21 launch |
| DeepSeek V4 Pro | $0.435 | $0.87 | Off-peak; 2x peak |
| GPT-5.6 Luna | $1 | $6 | Cheap OpenAI tier |
| Claude Haiku 4.5 | $1 | $5 | |
| Gemini 3.6 Flash | $1.50 | $7.50 | Google default |
| Claude Sonnet 5 | $2 | $10 | Intro → $3/$15 after Aug 31, 2026 |
| Grok 4.5 | $2 | $6 | Strong cheap frontier |
Cost Per Task (30K in / 5K out)
- DeepSeek V4 Pro (off-peak): ~$0.017
- GPT-5.6 Luna: ~$0.06
- Gemini 3.6 Flash: ~$0.0825
- Grok 4.5: ~$0.09
- Claude Sonnet 5 (intro): ~$0.11
- Claude Opus 5: ~$0.28
How To Choose
- Absolute cheapest, high volume → DeepSeek V4 Flash. Near-free cache hits make it ideal for RAG and agents.
- Cheap with more headroom → Grok 4.5 or Gemini 3.6 Flash — frontier-ish quality at a fraction of Opus/Sol.
- Cheapest OpenAI → GPT-5.6 Luna ($1/$6).
- Watch the clock → DeepSeek’s 2x peak-hour policy (Beijing windows) can double costs — schedule batch jobs off-peak.
Watch Outs
- Sonnet 5’s intro pricing ends Aug 31, 2026 → $3/$15 after.
- Prices move fast — always confirm on the vendor’s official pricing page before committing.
Verdict
- Cheapest overall → DeepSeek V4 Flash
- Best cheap frontier → Grok 4.5 / Gemini 3.6 Flash
- Cheapest OpenAI → GPT-5.6 Luna
Sources
- DeepSeek — pricing: api-docs.deepseek.com
- Google — Gemini API pricing: ai.google.dev/pricing
- Anthropic — pricing: anthropic.com/pricing