AI agents · OpenClaw · self-hosting · automation

Quick Answer

AI API Prices August 2026: Who Cut, Who Raised

Published:

The Short Answer

August 2026 broke the “inference always gets cheaper” narrative. Four major providers moved prices within eleven days, in two different directions.

  • OpenAI cut its flagship 20–33%
  • Anthropic cancelled a planned increase
  • Google launched at a half-price introductory rate
  • DeepSeek raised prices and added peak pricing

The Full August 2026 Picture

DateProviderMoveDetail
Jul 30OpenAI⬇️ CutTerra −20%, Luna −80%
~Aug 10Anthropic⬇️ Cancelled hikeSonnet 5 stays $2/$10, permanently
Aug 13Google⬇️ Intro rateGemini 3.7 Flash $0.75/$3.75 (50% off)
Aug 16DeepSeek⬆️ RaisedV4 Pro & Flash repriced, peak/off-peak split
Aug 21OpenAI⬇️ CutSol $5/$30 → $4/$20

Current Rates, Verified August 25, 2026

ModelInput / MTokOutput / MTok30K/5K taskStatus
Claude Opus 5$5.00$25.00~$0.275Standard
GPT-5.6 Sol$4.00$20.00~$0.22⏳ Promotional
GPT-5.6 Terra$2.00$12.00~$0.12Standard
Claude Sonnet 5$2.00$10.00~$0.11✅ Permanent
Grok 4.6$2.00$6.00~$0.09Standard
GLM-5.3$1.40$4.40~$0.064Standard
Gemini 3.7 Flash$0.75$3.75~$0.041⏳ Intro to Dec 31, 2026
DeepSeek V4 Pro$0.66$1.98~$0.030⬆️ Off-peak; 2x at peak
DeepSeek V4 Flash$0.22$0.66~$0.010⬆️ Off-peak; 2x at peak
GPT-5.6 Luna$0.20$1.20~$0.012Standard

The Frontier Order Flipped

The single most consequential change: GPT-5.6 Sol is now cheaper than Claude Opus 5, having been more expensive before August 21.

  • Sol: ~$0.22 per 30K-in/5K-out task
  • Opus 5: ~$0.275 per equivalent task

Sol’s output cut (33%) was steeper than its input cut (20%), which is deliberate. Agentic workloads are output-heavy — an agent that reads a codebase once then writes across dozens of turns generates far more output than input. Cutting output pricing hardest is a targeted bid for exactly the workload class where Anthropic has been strongest.

DeepSeek’s Peak Pricing Is the Underrated Story

DeepSeek repriced at 16:00 UTC on August 16, 2026, moving from flat rates to a peak/off-peak split where peak costs exactly double off-peak.

Peak hours: 01:00–04:00 and 06:00–10:00 UTC (7 hours). The remaining 17 hours are off-peak.

This matters geographically:

  • US and EU business hours land mostly in off-peak — you get the cheap rate by default
  • Asia-Pacific business hours land substantially in peak — you pay double

Cache pricing moved much harder than the headline rates. V4 Pro cache-hit went from $0.003625 to $0.022 off-peak — up to a 12x increase. Any architecture built around aggressive prompt caching on DeepSeek needs its cost model rebuilt, not adjusted.

Never quote a single flat DeepSeek price. Since August 16 that number does not exist.

What Is Temporary vs Permanent

This is the distinction that will cost teams money in Q4 2026 and Q1 2027.

Temporary — will revert:

  • GPT-5.6 Sol $4/$20 — reported as roughly three months
  • Gemini 3.7 Flash $0.75/$3.75 — reverts to $1.50/$7.50 on January 1, 2027, a 100% increase

Permanent — safe to build on:

  • Claude Sonnet 5 $2/$10 — Anthropic explicitly stated the scheduled September 1 increase to $3/$15 will not occur
  • OpenAI’s July 30 Terra and Luna cuts

The rule: budget at the post-promotional rate, spend at the promotional rate, keep the difference as margin. A team that priced its product on Gemini 3.7 Flash’s introductory rate is facing a doubling of that line on January 1 — a date that is knowable today.

What To Actually Do

  1. Re-run your cost model. If you moved workloads down-tier for budget reasons, the frontier tier may now be affordable for some of them.
  2. Tag every rate as standard or promotional in your cost tracking, with the revert date.
  3. Check your DeepSeek cache assumptions if you use it — the cache economics changed by an order of magnitude, not a percentage.
  4. Do not migrate on price alone. A 20% rate difference is dwarfed by output-token efficiency differences between models. Measure cost per completed task on your own workload before moving anything.

Sources