AI agents · OpenClaw · self-hosting · automation

Quick Answer

Cheapest AI API Aug 2026: V4 Flash vs Flash-Lite vs Luna

Published:

The Short Answer

For the cheapest capable AI API in August 2026, it’s a two-horse race on raw tokens: Gemini 2.5 Flash-Lite ($0.10/$0.40) and DeepSeek V4 Flash ($0.14/$0.28). GPT-5.6 Luna just cut 80% to $0.20/$1.20 — competitive, with OpenAI’s ecosystem as its edge, but not the absolute floor.

Side-by-Side (per MTok)

ModelInputOutputNotes
Gemini 2.5 Flash-Lite$0.10$0.40Cheapest input
DeepSeek V4 Flash$0.14$0.28Cheapest output; 2x during peak (1-4 & 6-10 UTC)
GPT-5.6 Luna$0.20$1.20Cut 80% Jul 30, 2026

Cost by Workload

A 30K-in / 5K-out task:

ModelCost
Gemini 2.5 Flash-Lite$0.005
DeepSeek V4 Flash (off-peak)~$0.0056
GPT-5.6 Luna$0.012

An output-heavy 10K-in / 30K-out task:

ModelCost
DeepSeek V4 Flash (off-peak)~$0.0098
Gemini 2.5 Flash-Lite$0.013
GPT-5.6 Luna$0.038

Output-heavy → DeepSeek V4 Flash wins. Input-heavy → Gemini 2.5 Flash-Lite wins.

How To Choose

  • Absolute cheapest, input-heavy (RAG, classification) → Gemini 2.5 Flash-Lite.
  • Absolute cheapest, output-heavy (generation, long answers) → DeepSeek V4 Flash — but watch the peak-hours 2x multiplier (1-4 & 6-10 UTC).
  • Want the OpenAI stack + tool use at low cost → GPT-5.6 Luna — pricier per token, but you get GPT-5.6 reliability and ecosystem.

The Trade-Offs

  • DeepSeek peak pricing: off-peak rates double during peak windows; schedule batch jobs off-peak to keep the low rate.
  • Data residency: DeepSeek is China-hosted — a factor for regulated data. Gemini and OpenAI offer Western-region hosting.
  • Ecosystem: Luna’s higher price buys OpenAI tooling, function calling maturity, and brand reliability.

Verdict

  • Cheapest input-heavy → Gemini 2.5 Flash-Lite
  • Cheapest output-heavy → DeepSeek V4 Flash
  • Cheap + OpenAI ecosystem → GPT-5.6 Luna

Sources