GPT-5.6 Luna Price Cut 80%: New Rates (Aug 2026)
The Short Answer
On July 30, 2026, OpenAI cut GPT-5.6 Luna by ~80% to $0.20 input / $1.20 output per million tokens (was $1/$6), and cut GPT-5.6 Terra by 20% to $2/$12 (was $2.50/$15). OpenAI confirmed both on its official price-performance page. It’s a defensive move against cheaper rivals.
New Pricing
| Model | New price (per MTok) | Old price | Change |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 / $1.20 | $1 / $6 | −80% |
| GPT-5.6 Terra | $2 / $12 | $2.50 / $15 | −20% |
| GPT-5.6 Sol | $5 / $30 | (unchanged) | — |
Effective July 30, 2026. Luna is OpenAI’s fastest, cheapest tier; Terra is the balanced everyday model; Sol is the flagship.
How It Stacks Up
| Model | Input / Output (per MTok) |
|---|---|
| Gemini 2.5 Flash-Lite | $0.10 / $0.40 |
| DeepSeek V4 Flash | $0.14 / $0.28 |
| GPT-5.6 Luna (new) | $0.20 / $1.20 |
| Grok 4.5 | $2 / $6 |
| Claude Haiku 4.5 | $1 / $5 |
On raw tokens, DeepSeek V4 Flash and Gemini 2.5 Flash-Lite are still cheaper than Luna. Luna’s advantage is the GPT-5.6 ecosystem, tool use, and brand reliability — not the absolute floor.
Why It Happened
The cut is a direct response to the AI price war. Cheap open-weight models (DeepSeek, Qwen, GLM) and Google’s Flash-Lite tier have made low-cost inference a commodity. As enterprises scrutinize token spend, OpenAI needed Luna competitive on high-volume jobs — classification, extraction, routing, summarization — where output volume dominates cost.
What to Do
- High-volume, cost-sensitive jobs → Luna is now viable inside the OpenAI stack; benchmark it against Gemini 2.5 Flash-Lite and DeepSeek V4 Flash.
- Absolute cheapest per token → DeepSeek V4 Flash or Gemini 2.5 Flash-Lite still win.
- Already on GPT-5.6 → the Luna cut may reshuffle your routing tiers; re-run your cost model.
Sources
- OpenAI — advancing the price-performance frontier with GPT-5.6: openai.com
- CNBC — OpenAI cuts prices for two GPT-5.6 models: cnbc.com
- VentureBeat — AI price wars, Luna cut 80%: venturebeat.com