AI agents · OpenClaw · self-hosting · automation

Quick Answer

DeepSeek V4-Flash vs Nova vs Qwen 3.7 Flash: Cheapest API

Published:

The Short Answer

For the cheapest AI API in August 2026: DeepSeek V4-Flash ($0.14/$0.28) for frontier-adjacent quality at rock-bottom cost, Amazon Nova Micro/Lite for the absolute lowest price on light tasks, and Qwen 3.7 Flash for cheap open-model coding.

The Comparison

DeepSeek V4-FlashAmazon Nova MicroQwen 3.7 Flash
Input (per MTok)$0.14$0.035$0.03
Output (per MTok)$0.28$0.14$0.13
Cache-hit input$0.0028
TierFrontier-adjacentLight/fastLight coding
Peak surcharge2x planned (date TBD)NoNo

Where Each Wins

  • DeepSeek V4-Flash → best value at the frontier. Widely called the “price floor of the frontier-adjacent API market” — roughly 10x to 90x cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). The cache-hit input price of $0.0028/MTok makes cache-heavy agent loops extraordinarily cheap.
  • Amazon Nova Micro/Lite → absolute floor for light work. At $0.035/$0.14 (Micro) and $0.06/$0.24 (Lite), Nova undercuts almost everything for classification, extraction, and short responses inside AWS.
  • Qwen 3.7 Flash → cheap coding. At $0.03/$0.13, it targets budget coding tasks where you want an open-weight lineage and low latency.

How To Choose

  1. Frontier-adjacent quality, lowest cost → DeepSeek V4-Flash.
  2. Highest-volume light tasks in AWS → Amazon Nova Micro/Lite.
  3. Cheap coding, open-model lineage → Qwen 3.7 Flash.

Note: for the absolute cheapest production API overall, Llama 3.1 8B Instruct runs about $0.02/MTok input — but it’s a small model, not frontier-adjacent.

Verdict

  • Best value at the frontier → DeepSeek V4-Flash
  • Cheapest for light workloads → Amazon Nova Micro
  • Cheapest budget coding → Qwen 3.7 Flash

Sources