AI API Prices August 2026: Who Cut, Who Raised
The Short Answer
August 2026 broke the “inference always gets cheaper” narrative. Four major providers moved prices within eleven days, in two different directions.
- OpenAI cut its flagship 20–33%
- Anthropic cancelled a planned increase
- Google launched at a half-price introductory rate
- DeepSeek raised prices and added peak pricing
The Full August 2026 Picture
| Date | Provider | Move | Detail |
|---|---|---|---|
| Jul 30 | OpenAI | ⬇️ Cut | Terra −20%, Luna −80% |
| ~Aug 10 | Anthropic | ⬇️ Cancelled hike | Sonnet 5 stays $2/$10, permanently |
| Aug 13 | ⬇️ Intro rate | Gemini 3.7 Flash $0.75/$3.75 (50% off) | |
| Aug 16 | DeepSeek | ⬆️ Raised | V4 Pro & Flash repriced, peak/off-peak split |
| Aug 21 | OpenAI | ⬇️ Cut | Sol $5/$30 → $4/$20 |
Current Rates, Verified August 25, 2026
| Model | Input / MTok | Output / MTok | 30K/5K task | Status |
|---|---|---|---|---|
| Claude Opus 5 | $5.00 | $25.00 | ~$0.275 | Standard |
| GPT-5.6 Sol | $4.00 | $20.00 | ~$0.22 | ⏳ Promotional |
| GPT-5.6 Terra | $2.00 | $12.00 | ~$0.12 | Standard |
| Claude Sonnet 5 | $2.00 | $10.00 | ~$0.11 | ✅ Permanent |
| Grok 4.6 | $2.00 | $6.00 | ~$0.09 | Standard |
| GLM-5.3 | $1.40 | $4.40 | ~$0.064 | Standard |
| Gemini 3.7 Flash | $0.75 | $3.75 | ~$0.041 | ⏳ Intro to Dec 31, 2026 |
| DeepSeek V4 Pro | $0.66 | $1.98 | ~$0.030 | ⬆️ Off-peak; 2x at peak |
| DeepSeek V4 Flash | $0.22 | $0.66 | ~$0.010 | ⬆️ Off-peak; 2x at peak |
| GPT-5.6 Luna | $0.20 | $1.20 | ~$0.012 | Standard |
The Frontier Order Flipped
The single most consequential change: GPT-5.6 Sol is now cheaper than Claude Opus 5, having been more expensive before August 21.
- Sol: ~$0.22 per 30K-in/5K-out task
- Opus 5: ~$0.275 per equivalent task
Sol’s output cut (33%) was steeper than its input cut (20%), which is deliberate. Agentic workloads are output-heavy — an agent that reads a codebase once then writes across dozens of turns generates far more output than input. Cutting output pricing hardest is a targeted bid for exactly the workload class where Anthropic has been strongest.
DeepSeek’s Peak Pricing Is the Underrated Story
DeepSeek repriced at 16:00 UTC on August 16, 2026, moving from flat rates to a peak/off-peak split where peak costs exactly double off-peak.
Peak hours: 01:00–04:00 and 06:00–10:00 UTC (7 hours). The remaining 17 hours are off-peak.
This matters geographically:
- US and EU business hours land mostly in off-peak — you get the cheap rate by default
- Asia-Pacific business hours land substantially in peak — you pay double
Cache pricing moved much harder than the headline rates. V4 Pro cache-hit went from $0.003625 to $0.022 off-peak — up to a 12x increase. Any architecture built around aggressive prompt caching on DeepSeek needs its cost model rebuilt, not adjusted.
Never quote a single flat DeepSeek price. Since August 16 that number does not exist.
What Is Temporary vs Permanent
This is the distinction that will cost teams money in Q4 2026 and Q1 2027.
Temporary — will revert:
- GPT-5.6 Sol $4/$20 — reported as roughly three months
- Gemini 3.7 Flash $0.75/$3.75 — reverts to $1.50/$7.50 on January 1, 2027, a 100% increase
Permanent — safe to build on:
- Claude Sonnet 5 $2/$10 — Anthropic explicitly stated the scheduled September 1 increase to $3/$15 will not occur
- OpenAI’s July 30 Terra and Luna cuts
The rule: budget at the post-promotional rate, spend at the promotional rate, keep the difference as margin. A team that priced its product on Gemini 3.7 Flash’s introductory rate is facing a doubling of that line on January 1 — a date that is knowable today.
What To Actually Do
- Re-run your cost model. If you moved workloads down-tier for budget reasons, the frontier tier may now be affordable for some of them.
- Tag every rate as standard or promotional in your cost tracking, with the revert date.
- Check your DeepSeek cache assumptions if you use it — the cache economics changed by an order of magnitude, not a percentage.
- Do not migrate on price alone. A 20% rate difference is dwarfed by output-token efficiency differences between models. Measure cost per completed task on your own workload before moving anything.