Claude Sonnet 5.5 vs GPT-6 Sol vs Grok 4.7: $2 Tier (2026)
The short answer
Three frontier-lab models now sit at $2 per million input tokens, and they are not interchangeable. Claude Sonnet 5.5 (September 28, 2026, $2/$10) has the best published coding and knowledge-work scores of the three. GPT-6 Sol (September 22, 2026, $2/$10) has the biggest context window (1,050,000 tokens) and the deepest OpenAI-stack integration. Grok 4.7 (September 21, 2026, $2/$6) has the cheapest output and an independent Artificial Analysis Intelligence Index of 46 — but it burns roughly twice the output tokens of its predecessor. For agentic coding and documents pick Sonnet 5.5; for very long prompts inside OpenAI pick Sol; for output-heavy generation where the $6 rate dominates pick Grok 4.7.
Side by side
| Claude Sonnet 5.5 | GPT-6 Sol | Grok 4.7 | |
|---|---|---|---|
| Vendor | Anthropic | OpenAI | xAI (SpaceXAI) |
| Released | Sep 28, 2026 | Sep 22, 2026 | Sep 21, 2026 |
| Input / output (per MTok) | $2 / $10 | $2 / $10 | $2 / $6 |
| Cached input | $0.20 | $0.20 | $0.50 |
| Cache write | $2.50 | $2.50 | — |
| Long-prompt surcharge | None | >272K: 2x input, 1.5x output | ≥200K: $4 / $1 cached / $12 |
| Context / max output | 1M / 128K | 1,050,000 (922K in) / 128K | 500K / — |
| Knowledge cutoff | June 2026 | Apr 20, 2026 | May 2026 |
| Reasoning control | Adaptive thinking, 5 effort levels | 6 effort levels (none → max) | Reasoning; Fast variant 2x price, 2x speed |
| Batch / flex | Batch 50% | Batch/Flex 50%, fast mode 2x | — |
| AA Intelligence Index | Not scored (Sep 29) | 48 | 46 |
| FrontierCode 1.1 | 52.1% (xhigh) | 49.3% | — |
| GDPval-AA v2.1 | 1844 | 1487 | — |
| AA-Briefcase v1.1 | 1811 | 1483 | — |
| Chartography (no tools) | 61.6% | 53.6% | — |
| Terminal-Bench 4.0 | 70.6% | not reported by OpenAI | 26% (AA) / 38.0% (xAI table) |
| Output tokens per AA task | not measured | — | ~81K (4.6: ~36K) |
| Reference 30K-in / 5K-out | $0.11 | $0.11 | $0.09 |
| In Cursor | Yes | Until Nov 12, 2026 (proposed cutoff) | Yes (flagship) |
Anthropic ran the GPT-6 Sol comparisons itself; OpenAI has not published Terminal-Bench for Sol. Grok 4.7’s Terminal-Bench figure is Artificial Analysis’s 26% versus xAI’s own 38.0%; the discrepancy is unresolved. GPT-6 Sol and Grok 4.7 details from what are GPT-6 Sol and Luna and what is Grok 4.7.
Price: identical headlines, different fine print
At list price Sonnet 5.5 and GPT-6 Sol are the same model economically: $2/$10, $0.20 cache reads, $2.50 cache writes, 50% batch. Grok 4.7 undercuts both on output ($6) but charges 2.5x more for cache reads ($0.50), which matters for agents that re-read a large fixed context every step.
The differences are at the edges:
- Long prompts. GPT-6 Sol reprices the entire request once input exceeds 272,000 tokens (2x input, 1.5x output). Grok 4.7 does the same at 200,000 tokens, jumping to $4/$12. Sonnet 5.5 has no such tier: a 600K-token context read costs $1.20 on Sonnet 5.5, $2.40 on GPT-6 Sol and $2.40 on Grok 4.7.
- Tokens per task. Artificial Analysis measured Grok 4.7 at ~81K output tokens per Intelligence Index task, about double Grok 4.6’s ~36K, so its per-task cost roughly doubled at the same list price. Anthropic claims Sonnet 5.5 needs “far fewer tokens” than Sonnet 5 (up to 30% cheaper per task). GPT-6 Sol was mainly a price story: 50% below GPT-5.6 Sol’s promo $4/$20, level on capability (AA 47 → 48). See how to audit an LLM provider for token inflation.
- Promo risk. GPT-5.6-era promo prices are scheduled to rise around November 2026; GPT-6 Sol’s $2/$10 is its launch list price, not a promo. Gemini 3.8 Flash ($0.75/$3.75) is the fourth option here but doubles to $1.50/$7.50 on January 1, 2027.
Capability: what the numbers can and cannot say
On the four benchmarks where Anthropic ran GPT-6 Sol alongside Sonnet 5.5, Sonnet 5.5 wins all four, and the GDPval-AA gap (1844 vs 1487) is large — it is the difference between “nearly Opus 5.5” and “Sonnet 5 territory” on real occupational tasks. Sonnet 5.5 also has the best computer-use score of Anthropic’s lineup below Opus (80.1% OSWorld 2.1).
Two cautions. First, these are vendor runs on a launch day; an independent Artificial Analysis score for Sonnet 5.5 did not exist as of September 29, 2026, whereas GPT-6 Sol (48) and Grok 4.7 (46) do have one. Second, the Sol-vs-Sonnet comparison skips xAI entirely because Anthropic did not include Grok in its table; Grok 4.7’s strengths in xAI’s own reporting are agentic coding and tool use, where it is Cursor’s flagship for long-running tasks.
Ecosystem and lock-in
- Sonnet 5.5 is live on the Claude Platform, AWS, Google Cloud and Azure with zero data retention, and it is the default in Claude Code and the Claude apps at Medium effort. It carries cyber and reasoning-extraction safeguards; higher-risk security prompts fall back to Sonnet 5.
- GPT-6 Sol is on the OpenAI API and in ChatGPT; it is the mid-tier of a lineup that also has GPT-6 Luna ($0.10/$0.50) and GPT-6 Astra ($10/$50). If you are in Cursor, note the proposed November 12, 2026 discontinuation of OpenAI models there.
- Grok 4.7 is in Cursor, Grok Build, GitHub Copilot and the Vercel AI Gateway, and it is the model SpaceXAI is wrapping in Nvidia’s new OpenShell agent runtime (see what is Nvidia Open Agent Safety Platform).
Decision rule
- Agentic coding, documents, slides, spreadsheets: Claude Sonnet 5.5 — best published scores, no long-context surcharge, cheapest cache reads (tied).
- Prompts routinely between 500K and 1M tokens: GPT-6 Sol has the largest window, but budget the >272K surcharge; Sonnet 5.5 is cheaper past 272K.
- Output-heavy generation (reports, code dumps) with short prompts: Grok 4.7’s $6 output wins on paper; verify tokens per task on your workload first.
- Need an independent benchmark today: GPT-6 Sol (AA 48) or Grok 4.7 (AA 46); Sonnet 5.5 is unscored as of September 29, 2026.
- Already on Cursor: Grok 4.7 or Sonnet 5.5; OpenAI models face a November 12 cutoff there.
Last verified: September 29, 2026. All prices are standard API list rates in USD per million tokens.