Claude Sonnet 5.5 vs Opus 5.5: Which to Use (Sep 2026)
The short answer
Claude Sonnet 5.5 (September 28, 2026) is half the price of Claude Opus 5.5 (September 22, 2026) and lands within two points of it on most of Anthropic’s published benchmarks, so it is the default for well-scoped coding, documents and chat; Opus 5.5 is still the model for open-ended, multi-hour work that needs sustained judgment. Sonnet 5.5 costs $2/$10 per million tokens versus $4/$20, generates 30%+ faster than Sonnet 5, and actually beats Opus 5.5 on Terminal-Bench 4.0 (70.6% vs 66.4%). Anthropic’s own line is that Opus 5.5 “remains clearly stronger at complex, open-ended work.”
Side by side
| Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|
| Released | Sep 28, 2026 | Sep 22, 2026 |
| Input / output (per MTok) | $2 / $10 | $4 / $20 |
| Cache read / 5m write / 1h write | $0.20 / $2.50 / — | $0.20 / $5 / $8 |
| Fast mode | — | $8 / $40 (API only) |
| Context / max output | 1M / 128K | 1M / 128K |
| Knowledge cutoff | June 2026 | June 2026 |
| Default effort (API) | High | Medium |
| Default effort (Claude Code, apps) | Medium | Medium |
| Effort levels | 5 (low → max) | 5 (low → max) |
| Cyber safeguards | Yes (falls back to Sonnet 5) | Yes |
| AA Intelligence Index | Not yet scored | 58 (current index) |
| Terminal-Bench 4.0 | 70.6% | 66.4% (xhigh) |
| FrontierCode 1.1 (Main) | 52.1% (xhigh) | 54.4% |
| CursorBench 4.0 | 55.5% | 57.8% |
| GDPval-AA v2.1 | 1844 | 1846 |
| AA-Briefcase v1.1 | 1811 | 1822 |
| Humanity’s Last Exam (tools) | 64.5% | 67.7% |
| OSWorld 2.1 (partial) | 80.1% | 81.8% |
| Chartography (no tools) | 61.6% | 64.4% |
| Reference 30K-in / 5K-out step | $0.11 | $0.22 |
Benchmarks are Anthropic’s, published September 28, 2026. Opus 5.5 specs and its AA score are from the Opus 5.5 launch.
Where Sonnet 5.5 is “good enough”
The two-point pattern is the story. On GDPval-AA (real deliverables across 44 occupations and nine industries) Sonnet 5.5 is at 1844 versus 1846. On computer use it is 80.1% versus 81.8%. On CursorBench, built from real Cursor sessions, it is 55.5% versus 57.8%. Those gaps are inside the noise of most production evals, and they come at half the token price plus a faster response.
Anthropic’s cost-vs-accuracy charts add a second argument: at Low or Medium effort, Sonnet 5.5 beats Sonnet 5’s best score for about a tenth of the cost per task, and on FrontierCode at High effort it scores 10 points above Sonnet 5 at roughly one fifteenth of the cost. If your workload was fine on Sonnet 5, Sonnet 5.5 is a free upgrade with a lower bill; if you were paying for Opus because Sonnet 5 fell short, re-run the eval before renewing that decision.
The Terminal-Bench inversion (70.6% vs 66.4%) is worth a careful read. Terminal-Bench 4.0 rewards multi-step command-line execution with tight scope, exactly the “well-scoped” work Anthropic aims Sonnet 5.5 at, and Sonnet 5.5’s habit of batching tool calls reduces step count. It does not mean Sonnet 5.5 is the better coding model overall — FrontierCode and CursorBench still go to Opus.
Where Opus 5.5 still earns 2x
- Open-ended, long-horizon work. Anthropic’s external testers and its own team found Opus 5.5 “clearly stronger” when the task needs sustained judgment: ambiguous specs, refactors across large systems, research synthesis. The Opus 5.5 migration notes cover its always-on thinking and default-medium effort.
- Hard reasoning. Humanity’s Last Exam with tools: 67.7% vs 64.5%.
- Max effort behaviour. Sonnet 5.5’s FrontierCode score drops from 52.1% at xhigh to 46.2% at max because it starts running multi-agent code review that overruns scope. Opus 5.5 at max effort emits roughly 119K output tokens per Artificial Analysis task, which is expensive but does not show that failure mode. If you need “throw everything at it,” Opus.
- An independent score. Opus 5.5 has an Artificial Analysis Intelligence Index of 58 (Fable 5.1 and GPT-6 Astra sit at 53). Sonnet 5.5 had no independent score as of September 29, 2026, so every number above is vendor-reported.
Cost per task, not per token
List price says 2x. Real spend depends on effort and tokens per task:
- At low/medium effort, Sonnet 5.5 is far cheaper than 2x: fewer thinking tokens and fewer tool calls. Balyasny measured ~121K tokens per answer on Sonnet 5.5 versus 497K on Sonnet 5; Base44 needed 3.6 build iterations versus 7.7 on Opus 5.
- At xhigh/max, Anthropic says Sonnet 5.5 “can perform comparably at a similar cost” to Opus 5.5 — the price advantage evaporates because Sonnet thinks longer to get there.
- Cache-heavy agents see identical $0.20 cache reads on both, so the 2x gap applies only to uncached input and to output.
Practical rule: run Sonnet 5.5 at medium first. Escalate to Opus 5.5 only for the tasks that fail, not by default. See how to choose an LLM reasoning effort level.
Mixing both in one system
Two API details decide the architecture. Sonnet 5.5 reads thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier, but not from Opus 5, Opus 5.5, Fable or Mythos; the API drops those blocks without an error. So a conversation that starts on Opus 5.5 and switches to Sonnet 5.5 loses its reasoning trail mid-stream. The supported pattern is the advisor tool: a Sonnet 5.5 executor can call Opus 5.5 (or Opus 5, Sonnet 5.5, Fable 5/5.1, Mythos 5/5.1) for advice, returned encrypted. Opus 4.8 and earlier are no longer accepted as advisors. Worker-Sonnet, reviewer-Opus is the cheapest configuration that keeps Opus judgment in the loop.
Decision rule
- Bug fixes, scoped features, tests, CLI tasks: Sonnet 5.5, medium effort.
- Docs, slides, spreadsheets, UI polish: Sonnet 5.5 — this is what Anthropic built it for.
- Chat products, latency-sensitive: Sonnet 5.5, low or medium; it is Anthropic’s fastest Sonnet.
- Multi-hour autonomous runs, ambiguous specs, architecture: Opus 5.5.
- Anything that needs an independent benchmark to justify: Opus 5.5, until Sonnet 5.5 is scored.
- Security research: both carry cyber safeguards; apply to the Cyber Verification Program.
Last verified: September 29, 2026. Prices are standard API list rates; see PRICING for the full table.