AI agents · OpenClaw · self-hosting · automation

Quick Answer

What Is Claude Sonnet 5.5? Anthropic's New Model (Sep 2026)

Published:

The short answer

Claude Sonnet 5.5 is Anthropic’s new mid-tier model, released September 28, 2026, at the same $2/$10 per million tokens as Sonnet 5 but with a step-change in agentic coding (70.6% on Terminal-Bench 4.0, up from 10.3%), 30%+ faster output and up to 30% lower cost per task. It is the second model in the Claude 5.5 family, six days behind Claude Opus 5.5, and it is positioned as the faster, cheaper complement: well-scoped everyday tasks, bug fixing, and polished documents, slides and spreadsheets. Claude Haiku 5.5 is next, “in the coming weeks.”

Specs at a glance

Claude Sonnet 5.5Claude Sonnet 5Claude Opus 5.5
ReleasedSep 28, 2026Jun 30, 2026Sep 22, 2026
Input / output (per MTok)$2 / $10$2 / $10$4 / $20
Cache read / 5m write$0.20 / $2.50$0.20 / $2.50$0.20 / $5
Context / max output1M / 128K1M / 128K1M / 128K
Knowledge cutoffJune 2026—June 2026
ThinkingAdaptive, on by defaultAdaptiveAdaptive, always on
Effort levelslow, medium, high, xhigh, max45
Default effortHigh (API), Medium (apps, Claude Code)HighMedium
Output speed vs Sonnet 530%+ faster——
Min cacheable prompt512 tokens1,024512
Model IDclaude-sonnet-5-5claude-sonnet-5claude-opus-5-5
CloudsClaude Platform, AWS, Google Cloud, Azuresamesame

The pricing line matters more than it looks: Sonnet 5.5 is now exactly half of Opus 5.5 on input, output and cache writes, and it matches GPT-6 Sol’s $2/$10 (see Sonnet 5.5 vs GPT-6 Sol vs Grok 4.7).

Benchmarks (vendor-reported, September 28, 2026)

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.0 (agentic coding)70.6%10.3%66.4% (xhigh)not reported
FrontierCode 1.1 (Main)52.1% xhigh / 46.2% max42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%—
GDPval-AA v2.1 (knowledge work)1844144918461487
AA-Briefcase v1.11811135918221483
Humanity’s Last Exam (with tools)64.5%54.9%67.7%—
OSWorld 2.1 (computer use, partial)80.1%57.0%81.8%—
Chartography (no tools)61.6%15.6%64.4%53.6%

Three things stand out. First, the Terminal-Bench jump from 10.3% to 70.6% is the largest single-generation move Anthropic has published for a Sonnet, and it puts Sonnet 5.5 above Opus 5.5 on that one test. Second, on GDPval-AA (real tasks across 44 occupations) Sonnet 5.5 is two points behind Opus 5.5 and about 400 points above Sonnet 5. Third, the FrontierCode “max” score (46.2%) is lower than “xhigh” (52.1%): at max effort the model more often spun up multi-agent code review, which caused timeouts and out-of-scope edits that FrontierCode penalizes. Anthropic’s own framing is that Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.” GPT-6 Sol numbers are Anthropic’s runs; OpenAI did not publish Terminal-Bench for Sol.

What changed versus Sonnet 5

  • Speed and efficiency. 30%+ faster generation, and fewer tokens and tool calls per task. Early testers saw it batch tool calls more aggressively than Sonnet 5. Customer data in the launch post: Balyasny Asset Management ~121K tokens per answer versus 497K on Sonnet 5; Base44 3.6 iterations per app build where Opus 5 took 7.7; Slack ~14% fewer output tokens with better results on almost all Slackbot evals; Zendesk 20% faster ticket handling.
  • Vision and long-horizon work. First Sonnet to beat Pokémon Red from screenshots alone; Chartography jumps from 15.6% to 61.6%.
  • Writing and design. Clearer prose than the previous generation; testers noted it follows slide templates well enough that a 10-slide earnings review draft was judged “ready to send as is” by two experts in an internal test.
  • API surface. thinking: {"type": "disabled"} is gone (use between_tools), forced tool_choice of any/tool returns a 400, thinking blocks are signed over the conversation for accounts created on or after August 31, 2026, computer use moves to computer_toolset_20260801 on the Claude API and Google Cloud, and the minimum cacheable prompt drops to 512 tokens. Full checklist in how to migrate to Claude Sonnet 5.5.

Safety and safeguards

Because Sonnet 5.5’s cyber capabilities are comparable to Opus 5’s, it is the first Sonnet to ship with real-time cyber safeguards: routine bug finding and fixing works, but higher-risk cybersecurity requests visibly fall back to Sonnet 5. Anthropic’s expanded Cyber Verification Program will offer tiered access on Sonnet 5.5, Opus 5.5 and Mythos models. Biology safeguards are unchanged from Sonnet 5. It is also the first Sonnet with reasoning-extraction classifiers (to blunt distillation attacks, the subject of Anthropic’s September 2026 threat report), and it expands preserved thinking so that thinking blocks cannot be moved between accounts. On Anthropic’s ~1,850-scenario behavioral audit it matches or improves on Sonnet 5 on most alignment measures, and on containment evaluations it is the least likely of any Claude model to probe the limits of its sandbox. A refusal now returns stop_reason: "refusal" with a category (cyber, bio, frontier_llm, reasoning_extraction, general_harms).

Who should use it

  • Agentic coding on a budget: Sonnet 5.5 at Medium effort beats Sonnet 5’s best Terminal-Bench score for less than a tenth of the cost per task. Start at medium for well-specified tasks and high for longer ones.
  • Documents, slides, spreadsheets, UI polish: this is where Anthropic explicitly points it.
  • Chat and latency-sensitive products: low or medium effort; 30%+ faster than Sonnet 5.
  • Not for: open-ended, multi-hour judgment work — that is still Opus 5.5 territory. Decision rules in Sonnet 5.5 vs Opus 5.5.

Last verified: September 29, 2026. All benchmark figures are Anthropic’s own; no independent Artificial Analysis index score for Sonnet 5.5 was published as of this date.

Sources