Quick Answer
DeepSeek V4-Flash vs Nova vs Qwen 3.7 Flash: Cheapest API
The Short Answer
For the cheapest AI API in August 2026: DeepSeek V4-Flash ($0.14/$0.28) for frontier-adjacent quality at rock-bottom cost, Amazon Nova Micro/Lite for the absolute lowest price on light tasks, and Qwen 3.7 Flash for cheap open-model coding.
The Comparison
| DeepSeek V4-Flash | Amazon Nova Micro | Qwen 3.7 Flash | |
|---|---|---|---|
| Input (per MTok) | $0.14 | $0.035 | $0.03 |
| Output (per MTok) | $0.28 | $0.14 | $0.13 |
| Cache-hit input | $0.0028 | — | — |
| Tier | Frontier-adjacent | Light/fast | Light coding |
| Peak surcharge | 2x planned (date TBD) | No | No |
Where Each Wins
- DeepSeek V4-Flash → best value at the frontier. Widely called the “price floor of the frontier-adjacent API market” — roughly 10x to 90x cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). The cache-hit input price of $0.0028/MTok makes cache-heavy agent loops extraordinarily cheap.
- Amazon Nova Micro/Lite → absolute floor for light work. At $0.035/$0.14 (Micro) and $0.06/$0.24 (Lite), Nova undercuts almost everything for classification, extraction, and short responses inside AWS.
- Qwen 3.7 Flash → cheap coding. At $0.03/$0.13, it targets budget coding tasks where you want an open-weight lineage and low latency.
How To Choose
- Frontier-adjacent quality, lowest cost → DeepSeek V4-Flash.
- Highest-volume light tasks in AWS → Amazon Nova Micro/Lite.
- Cheap coding, open-model lineage → Qwen 3.7 Flash.
Note: for the absolute cheapest production API overall, Llama 3.1 8B Instruct runs about $0.02/MTok input — but it’s a small model, not frontier-adjacent.
Verdict
- Best value at the frontier → DeepSeek V4-Flash
- Cheapest for light workloads → Amazon Nova Micro
- Cheapest budget coding → Qwen 3.7 Flash
Sources
- DeepSeek — pricing: deepseek.ai/pricing
- Caixin Global — DeepSeek V4-Flash release (Aug 1, 2026): caixinglobal.com
- AWS — Amazon Nova pricing: aws.amazon.com/nova