Claude Haiku 5.5 vs GPT-6 Luna vs Gemini 3.8 Flash (Oct 2026)
The short answer
Claude Haiku 5.5 (October 7, 2026) is now the best small model for prompts under 100,000 tokens: same $0.10/$0.50 price as GPT-6 Luna, higher scores on Anthropic’s launch benchmarks. Above 100K tokens, GPT-6 Luna is much cheaper. Gemini 3.8 Flash sits a tier higher in price and is the pick for audio and video input. Prices are list USD per million tokens from vendor pricing pages, read October 8, 2026; see current API prices for every model.
The comparison
| Claude Haiku 5.5 | GPT-6 Luna | Gemini 3.8 Flash | |
|---|---|---|---|
| Released | Oct 7, 2026 | Sep 22, 2026 | Sep 2, 2026 |
| Input / output (short prompts) | $0.10 / $0.50 (≤100K) | $0.10 / $0.50 (≤272K) | $0.75 / $3.75 (intro, through Dec 31, 2026) |
| Long-prompt price | $0.50 / $2.50 (>100K) | $0.20 / $0.75 (>272K) | Same rate; $1.50 / $7.50 from Jan 1, 2027 |
| Cached input | $0.01 (≤100K) | $0.01 | See Google pricing |
| Batch | 50% off ($0.05 / $0.25) | 50% off | 50% off |
| Context / max output | 1M / 128K | 1.05M / 128K | 1M / 64K |
| Reasoning | Adaptive, effort levels, default medium | Effort levels | Thinking levels |
| Input modalities | Text, image | Text, image | Text, image, audio, video |
Benchmarks (Anthropic’s launch table)
| Benchmark | Haiku 5.5 | GPT-6 Luna | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 16.4% | 0.0% | 70.6% |
| OSWorld 2.1 (offline subset) | 72.4% | 48.9% | 15.7% | 83.9% |
| FrontierCode 1.1 (Main) | 46.4% | 42.4% | — | 52.1% (xhigh) |
| GDPval-AA v2.1 | 1620 | 1437 | 735 | 1840 |
| Chartography (no tools) | 46.4% | 29.1% | 6.4% | 61.6% |
These are vendor-reported. No independent Artificial Analysis Intelligence Index score for Haiku 5.5 had been published when we checked on October 8, 2026.
What changed with Haiku 5.5
- Price. Input fell from $1 to $0.10 and output from $5 to $0.50 per million tokens for prompts up to 100K — 90% off list. Anthropic says prompts under 100K make up about 90% of requests to the previous Haiku and that Haiku 5.5 costs about 75% less to run on average.
- A hidden cost. The new tokenizer counts more tokens for the same text; Simon Willison measured about 1.25x versus Haiku 4.5. Reasoning cannot be switched off and defaults to medium effort, so output tokens are higher than a non-reasoning model’s.
- The 100K cliff. A 101K-token prompt pays five times the rate for the whole request. Keep RAG contexts and conversation history under the line, or route long prompts to Luna.
- Alongside it, Anthropic halved Sonnet 5.5 cache reads to $0.10 per million (about 20% cheaper on agentic work) and is adding monthly API credits for Max ($100 or $200) and Team (up to $500) subscribers.
Which to pick
- High-volume classification, extraction, summaries, support under 100K tokens: Haiku 5.5.
- Long documents or long agent histories above 100K tokens: GPT-6 Luna.
- Audio or video input, Google Cloud stack: Gemini 3.8 Flash — and budget for the January 2027 doubling.
- Subagents under Opus 5.5 or Sonnet 5.5: Haiku 5.5, which Anthropic built for that role.
Test on your own prompts at the effort level you will run: per-task cost depends on tokens consumed, not just list price — see token efficiency vs token price. Earlier budget-tier comparison: GPT-6 Luna vs Gemini 3.8 Flash vs DeepSeek V4.1 Flash.
Last verified: October 8, 2026.