AI agents · OpenClaw · self-hosting · automation

Quick Answer

GPT-5.6 Luna Price Cut 80%: New Rates (Aug 2026)

Published:

The Short Answer

On July 30, 2026, OpenAI cut GPT-5.6 Luna by ~80% to $0.20 input / $1.20 output per million tokens (was $1/$6), and cut GPT-5.6 Terra by 20% to $2/$12 (was $2.50/$15). OpenAI confirmed both on its official price-performance page. It’s a defensive move against cheaper rivals.

New Pricing

ModelNew price (per MTok)Old priceChange
GPT-5.6 Luna$0.20 / $1.20$1 / $6−80%
GPT-5.6 Terra$2 / $12$2.50 / $15−20%
GPT-5.6 Sol$5 / $30(unchanged)

Effective July 30, 2026. Luna is OpenAI’s fastest, cheapest tier; Terra is the balanced everyday model; Sol is the flagship.

How It Stacks Up

ModelInput / Output (per MTok)
Gemini 2.5 Flash-Lite$0.10 / $0.40
DeepSeek V4 Flash$0.14 / $0.28
GPT-5.6 Luna (new)$0.20 / $1.20
Grok 4.5$2 / $6
Claude Haiku 4.5$1 / $5

On raw tokens, DeepSeek V4 Flash and Gemini 2.5 Flash-Lite are still cheaper than Luna. Luna’s advantage is the GPT-5.6 ecosystem, tool use, and brand reliability — not the absolute floor.

Why It Happened

The cut is a direct response to the AI price war. Cheap open-weight models (DeepSeek, Qwen, GLM) and Google’s Flash-Lite tier have made low-cost inference a commodity. As enterprises scrutinize token spend, OpenAI needed Luna competitive on high-volume jobs — classification, extraction, routing, summarization — where output volume dominates cost.

What to Do

  • High-volume, cost-sensitive jobs → Luna is now viable inside the OpenAI stack; benchmark it against Gemini 2.5 Flash-Lite and DeepSeek V4 Flash.
  • Absolute cheapest per token → DeepSeek V4 Flash or Gemini 2.5 Flash-Lite still win.
  • Already on GPT-5.6 → the Luna cut may reshuffle your routing tiers; re-run your cost model.

Sources

  • OpenAI — advancing the price-performance frontier with GPT-5.6: openai.com
  • CNBC — OpenAI cuts prices for two GPT-5.6 models: cnbc.com
  • VentureBeat — AI price wars, Luna cut 80%: venturebeat.com