AI agents · OpenClaw · self-hosting · automation

Quick Answer

Claude Haiku 5.5 vs GPT-6 Luna vs Gemini 3.8 Flash (Oct 2026)

Published:

The short answer

Claude Haiku 5.5 (October 7, 2026) is now the best small model for prompts under 100,000 tokens: same $0.10/$0.50 price as GPT-6 Luna, higher scores on Anthropic’s launch benchmarks. Above 100K tokens, GPT-6 Luna is much cheaper. Gemini 3.8 Flash sits a tier higher in price and is the pick for audio and video input. Prices are list USD per million tokens from vendor pricing pages, read October 8, 2026; see current API prices for every model.

The comparison

Claude Haiku 5.5GPT-6 LunaGemini 3.8 Flash
ReleasedOct 7, 2026Sep 22, 2026Sep 2, 2026
Input / output (short prompts)$0.10 / $0.50 (≤100K)$0.10 / $0.50 (≤272K)$0.75 / $3.75 (intro, through Dec 31, 2026)
Long-prompt price$0.50 / $2.50 (>100K)$0.20 / $0.75 (>272K)Same rate; $1.50 / $7.50 from Jan 1, 2027
Cached input$0.01 (≤100K)$0.01See Google pricing
Batch50% off ($0.05 / $0.25)50% off50% off
Context / max output1M / 128K1.05M / 128K1M / 64K
ReasoningAdaptive, effort levels, default mediumEffort levelsThinking levels
Input modalitiesText, imageText, imageText, image, audio, video

Benchmarks (Anthropic’s launch table)

BenchmarkHaiku 5.5GPT-6 LunaHaiku 4.5Sonnet 5.5
Terminal-Bench 4.039.2%16.4%0.0%70.6%
OSWorld 2.1 (offline subset)72.4%48.9%15.7%83.9%
FrontierCode 1.1 (Main)46.4%42.4%—52.1% (xhigh)
GDPval-AA v2.1162014377351840
Chartography (no tools)46.4%29.1%6.4%61.6%

These are vendor-reported. No independent Artificial Analysis Intelligence Index score for Haiku 5.5 had been published when we checked on October 8, 2026.

What changed with Haiku 5.5

  • Price. Input fell from $1 to $0.10 and output from $5 to $0.50 per million tokens for prompts up to 100K — 90% off list. Anthropic says prompts under 100K make up about 90% of requests to the previous Haiku and that Haiku 5.5 costs about 75% less to run on average.
  • A hidden cost. The new tokenizer counts more tokens for the same text; Simon Willison measured about 1.25x versus Haiku 4.5. Reasoning cannot be switched off and defaults to medium effort, so output tokens are higher than a non-reasoning model’s.
  • The 100K cliff. A 101K-token prompt pays five times the rate for the whole request. Keep RAG contexts and conversation history under the line, or route long prompts to Luna.
  • Alongside it, Anthropic halved Sonnet 5.5 cache reads to $0.10 per million (about 20% cheaper on agentic work) and is adding monthly API credits for Max ($100 or $200) and Team (up to $500) subscribers.

Which to pick

  1. High-volume classification, extraction, summaries, support under 100K tokens: Haiku 5.5.
  2. Long documents or long agent histories above 100K tokens: GPT-6 Luna.
  3. Audio or video input, Google Cloud stack: Gemini 3.8 Flash — and budget for the January 2027 doubling.
  4. Subagents under Opus 5.5 or Sonnet 5.5: Haiku 5.5, which Anthropic built for that role.

Test on your own prompts at the effort level you will run: per-task cost depends on tokens consumed, not just list price — see token efficiency vs token price. Earlier budget-tier comparison: GPT-6 Luna vs Gemini 3.8 Flash vs DeepSeek V4.1 Flash.

Last verified: October 8, 2026.

Sources