AI agents · OpenClaw · self-hosting · automation

Quick Answer

Claude Opus 5.5 vs Grok 4.7 vs MiMo-V2.6-Pro (Sep 2026)

Published:

The short answer

Claude Opus 5.5 is the best model of the three and the most expensive; Grok 4.7 and MiMo-V2.6-Pro tie on intelligence, and MiMo costs a fifth as much. xAI shipped Grok 4.7 and Xiaomi shipped MiMo-V2.6-Pro on September 21, 2026, both landing at 46 on Artificial Analysis’s current Intelligence Index. Anthropic answered a day later with Opus 5.5 at 58, a 12-point gap, priced at $4/$20 per million tokens with $0.20 cache reads. Grok 4.7 at $2/$6 and MiMo-V2.6-Pro at $0.435/$0.87 make the mid-frontier the most crowded tier in the market, and the open-weight entrant is the value pick unless you need xAI’s tooling.

Side by side

Claude Opus 5.5Grok 4.7MiMo-V2.6-Pro
ReleasedSeptember 22, 2026September 21, 2026September 21, 2026
VendorAnthropicxAIXiaomi
Input / output (per MTok)$4 / $20$2 / $6 (<200K prompt)$0.435 / $0.87
Cached input$0.20$0.50$0.0036
Long-prompt repricingNone≥200K tokens: $4 / $1 cached / $12 for whole requestNone stated
Fast mode$8 / $40, up to 2.5x speedFast variant: 2x price, 2x speedPro-UltraSpeed: up to 20x speed
Context / max output1M / 128K500K / not stated1M / 128K
Knowledge cutoffJune 2026May 2026Not stated
Modalities inText, imageText, imageText, image, audio, video
WeightsClosedClosedOpen (MIT), 1.02T MoE / 42B active
AA Intelligence Index584646
AA output tokens per task~119K (max effort)~81KNot published
Terminal-Bench 4.066.4% (Anthropic)26% (AA); 38.0% (xAI)Not published
OSWorld 2.081.8% partialNot publishedNot published
GDPval-AA v2.11846 EloNot publishedNot published
Clouds / hostsAnthropic, AWS, GCP, AzurexAI, Cursor, GitHub Copilot, Grok Build, Vercel AI GatewayXiaomi API, self-host, third-party inference
SafeguardsFable-class cyber/bio; Life Sciences and Cyber Verification ProgramsStandardStandard

Cost per task, worked

On the reference 30K-input / 5K-output task:

ModelFormulaCost
Claude Opus 5.530 × $0.004 + 5 × $0.020$0.22
Grok 4.730 × $0.002 + 5 × $0.006$0.09
MiMo-V2.6-Pro30 × $0.000435 + 5 × $0.00087$0.017

List price understates two things. Opus 5.5 at max effort emits about 119,000 output tokens per Artificial Analysis task, so the per-task cost at max is roughly flat with Opus 5 despite the 20% price cut; at its default medium effort Anthropic says typical workloads cost 40% less than Opus 5. Grok 4.7 emits about 81,000 output tokens per task, double Grok 4.6, so its per-task cost roughly doubled at unchanged list prices. MiMo’s per-task token counts are not yet published by AA; assume the list-price gap narrows somewhat in practice, but not by 5x.

Where each one wins

Claude Opus 5.5 wins anything agentic. Its 66.4% on Terminal-Bench 4.0 is 14 points over Opus 5 and 8.5 over GPT-6 Astra; its 81.8% on OSWorld 2.0 is the highest published computer-use score; 1846 Elo on GDPval-AA v2.1 leads Fable 5.1 (1735). Anthropic says it matches Astra on Terminal-Bench at about 40% of the cost. It is also the first model shipped since Anthropic’s “pace the frontier” essay, was tested pre-release by METR and Frontier Design, and ships with Fable-class safeguards, meaning cyber and biology tasks can fall back to Opus 4.8 or Opus 5 when the classifier intervenes. Breaking changes for migrators: thinking cannot be disabled, forced tool use is gone, and the computer_20251124 tool is rejected. See How to migrate to Claude Opus 5.5.

Grok 4.7 wins on distribution and long context without a mode switch. It is live in Cursor, GitHub Copilot, Grok Build and Vercel AI Gateway on day one, has a 500K-token window, and a new larger base model that xAI says improves agentic coding (xAI’s own Terminal-Bench 4.0 figure is 38.0%; Artificial Analysis measured 26%). The rate card is identical to Grok 4.6, but the token count per task doubled, and the ≥200K repricing to $4/$12 makes long-context work cost as much as Opus 5.5. Upgrade analysis in Grok 4.7 vs Grok 4.6.

MiMo-V2.6-Pro wins on price, openness and modality. It is the top open-weight model on the AA index (46, above Kimi K3, Qwen3.8 Max, DeepSeek V4.1 Flash at 39 and V4 Pro at 36), takes audio and video input, and costs $0.435/$0.87 with a $0.0036 cache hit, the same price as V2.5. MIT licensing means you can post-train it the way Harvey post-trained Kimi K3, and Xiaomi is shipping RL tooling alongside it for exactly that. The caveat is independent benchmarking: as of September 24, 2026, the agentic coding and computer-use numbers Anthropic and xAI publish have no MiMo equivalent yet. Details in What is Xiaomi MiMo-V2.6.

Decision rule

  • Open-ended coding, computer use, research, anything where the task fails if the model is not smart enough: Claude Opus 5.5.
  • You already build in Cursor, Copilot or Vercel and want a cheaper frontier-adjacent model inside those tools: Grok 4.7.
  • High-volume work, multimodal input, self-hosting, or the first hop of a router that escalates hard cases: MiMo-V2.6-Pro.
  • Grok 4.7 vs MiMo-V2.6-Pro head to head: same intelligence score, 5x price difference, and MiMo has the weights. Grok needs its ecosystem to justify the premium.

The week’s pattern, across GPT-6 Sol and Luna, Grok 4.7, MiMo-V2.6 and Opus 5.5, is that the frontier split into a capability tier (Opus 5.5, Fable 5.1, Astra) and a price tier (everything else at 46–48 on the index), with the price tier now including an MIT-licensed model. Build the router; do not pick one.

Last verified: September 24, 2026.

Sources