AI agents · OpenClaw · self-hosting · automation

Quick Answer

Claude 5.5 Family: Opus vs Sonnet vs Haiku Tier Guide

Published:

The short answer

Use Claude Sonnet 5.5 by default, Opus 5.5 for the hardest tasks, and Haiku 4.5 for cheap volume — because Haiku 5.5 does not exist yet. As of October 3, 2026, Anthropic has shipped two of three 5.5-tier models: Opus 5.5 on September 22 at $4/$20 per million tokens, and Sonnet 5.5 on September 28 at $2/$10. Haiku 5.5 was promised “in the coming weeks” and has not landed, leaving Haiku 4.5 (October 15, 2025, $1/$5) as the current small tier. Facts verified October 3, 2026.

The three tiers

Opus 5.5Sonnet 5.5Haiku 5.5Haiku 4.5 (current)
Status✅ Shipped✅ Shipped❌ Unreleased✅ Shipped
ReleasedSep 22, 2026Sep 28, 2026”coming weeks”Oct 15, 2025
Input / output$4 / $20$2 / $10Unannounced$1 / $5
vs predecessor~40% cheaper to run than Opus 5Same price as Sonnet 5, ~30% fewer tokens/task——
PositioningFlagship: demanding reasoning, long-horizon agentsAgentic coding workhorseHigh-volume, cost-sensitiveHigh-volume, cost-sensitive
Terminal-Bench 4.0—70.6% (Anthropic) / 64% (AA)——
AA Intelligence Index5856—17 (reasoning)
Output speed (AA)—~139 tok/s—~91 tok/s
AA cost per task—$0.59 (reasoning)—$0.28 (reasoning)
Prompt caching✅✅—✅ up to 90% saving
Batch discount✅✅—✅ 50%

Anthropic’s and Artificial Analysis’s Terminal-Bench figures for Sonnet 5.5 differ (70.6% vs 64%) — the usual harness gap. Use the independent number when comparing across vendors.

Sonnet 5.5 is the story

The interesting thing about this generation is that Sonnet 5.5 got cheaper per task while staying the same price per token. It kept Sonnet 5’s $2/$10 list rate, and Anthropic says it typically needs fewer tokens — potentially up to 30% less per task.

That cuts against the dominant 2026 pattern. Most upgrades this year held list price while quietly increasing reasoning-token burn, so your bill went up without any price change you could point at. Grok 4.7 is the clean example: same per-token price as 4.6, roughly 81,000 output tokens per task at xhigh effort against ~36,000 before. Anthropic going the other direction is worth noticing — and worth verifying on your own traffic, because per-task token use swings hard with effort settings.

The capability jump is also real: Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 (Anthropic’s harness) against Sonnet 5’s much lower baseline makes it genuinely competitive for unattended terminal agents, which was the previous tier’s weak spot. At ~139 tokens/second it is also fast enough for interactive coding, where Opus-class latency becomes irritating.

When Opus 5.5 earns its 2×

Opus 5.5 leads Anthropic’s lineup on the Artificial Analysis Intelligence Index (58 vs Sonnet 5.5’s 56) and is positioned for demanding reasoning, coding and long-horizon agentic work. The $4/$20 rate is 40% cheaper to run than Opus 5 for typical workloads, which is a meaningful cut at the flagship tier.

But a 2-point index gap is not a 2× price gap. Route to Opus when you can name the failure: multi-hour agent trajectories that lose coherence on Sonnet, architectural refactors across many files, research synthesis where a subtle reasoning error is expensive. Measure it. The cheapest reliable pattern in 2026 is Sonnet-first with escalation to Opus on failure or low confidence, not Opus-by-default.

The Haiku 5.5 gap, and what to do about it

Anthropic flagged Haiku 5.5 as part of the 5.5 family on both September 22 and September 28, 2026, saying it would arrive “in the coming weeks.” As of October 3, 2026 it has not — no model, no price, no benchmarks.

So the small tier is still Haiku 4.5 at $1/$5, released October 15, 2025, which is now a year old. It remains genuinely good at what it is for: Artificial Analysis puts it at $0.28 per task in reasoning mode against Sonnet 5.5’s $0.59, with ~91 tokens/second output and ~0.59s time-to-first-token in non-reasoning mode. Its real jobs are classification and content moderation at scale, data extraction and transformation, agent routing and control flow, coding sub-agents inside a multi-agent system, and powering free tiers or real-time chat where sub-second latency matters more than depth.

Practical advice: build on Haiku 4.5 now and keep the model id in config. When 5.5 ships, you swap a string and re-run your evals. Do not architect around an unreleased model — and note that an intelligence index of 17 (reasoning) means Haiku 4.5 is a genuinely small model, so validate it on your task rather than assuming it degrades gracefully from Sonnet.

Routing rules that work

  • Agentic coding, terminal agents, most production work → Sonnet 5.5. Default here.
  • Long-horizon agents, hardest refactors, high-stakes reasoning → Opus 5.5, on escalation rather than by default.
  • Classification, extraction, routing, moderation, free tiers → Haiku 4.5 today; re-evaluate when 5.5 lands.
  • Multi-agent systems → Opus or Sonnet as orchestrator, Haiku as sub-agents. The cost asymmetry ($0.28 vs $0.59 per task) is the whole reason this pattern pays.
  • Every tier → turn on prompt caching (up to 90% saving on cached input) and batch processing (50%) before you consider downgrading a tier. Those two levers usually beat tier changes and cost no accuracy.

Related: Gemini 4 Argon vs Sonnet 5.5 vs GPT-6.1 Sol, Claude Sonnet 5.5 vs GPT-6.1 Sol, Grok 4.8 vs Grok 4.7 vs Grok 5, what is Claude Frontier Academy.

Last verified: October 3, 2026. Prices from Anthropic; independent scores from Artificial Analysis.

Sources