AI agents · OpenClaw · self-hosting · automation

Quick Answer

Grok 4.8 vs Grok 4.7 vs Grok 5: xAI's Roadmap

Published:

The short answer

Grok 4.7 is the only one of these you can use: released September 21, 2026, $2/$6 per million tokens, 500K context. Grok 4.8 exists as a claim — Musk described a 2.5-trillion-parameter model on a from-scratch C++ training stack around September 13–14, 2026 — with no xAI model page, price or release note as of October 3, 2026. Grok 5 is in training, unscheduled, sits behind 4.8 and 4.9, and is the model Musk says he expects to reach AGI. Build on 4.7. Facts verified October 3, 2026.

Side by side

Grok 4.7Grok 4.8Grok 5
Status✅ Generally available❌ Announced only❌ In training
DateReleased Sep 21, 2026Described Sep 13–14, 2026No date
ParametersNot disclosed (4.7 reported ~2.1T)2.5T (Musk)Unknown
Input / output$2 / $6 per MTokUnannouncedUnannounced
Fast variant$4 / $12 (2× speed)——
Cached input$0.50 per MTok——
Context500KUnknownUnknown
Reasoning levelslow / medium / high / xhigh——
Model IDgrok-4.7——
Training stackExistingNew C++ stack, from scratchUnknown
WherexAI API, Grok Build, Cursor, GitHub Copilot, routersTrace in Cursor’s model list—

Note the pricing fine print on 4.7: requests over 200,000 tokens double the rate across all tokens in the request, the US regional endpoint costs 10% more, priority processing bills at 2×, and web search is $5 per 1,000 calls. There is no free API tier, but Grok Build lets you try it at no cost.

Grok 4.7: where it actually wins

Grok 4.7’s independent numbers are uneven, and the pattern is worth knowing before you pick it.

It is strong on long-running and domain work. On Artificial Analysis’s Briefcase v1.1 (multi-hour office tasks) it scores 1,657, ahead of GPT-5.6 Sol at 1,487 and just behind Claude Fable 5.1 at 1,678. On CursorBench 4.0 it posts 46.3%, beating GPT-5.6 Sol’s 41.7% and its own predecessor’s 40.4%, trailing Fable 5.1’s 51.8%. On EEBench (electrical engineering) it jumps to 64.0% from 53.0% for Grok 4.6, ahead of both Fable 5.1 (56.4%) and GPT-5.6 Sol (39.4%). On the Harvey legal agent benchmark it leads its listed competitors at 19.6% against GPT-5.6 Sol’s 2.5%.

It is weaker than xAI implies on terminal agents. xAI’s own harness reports 38.0% on Terminal-Bench 4.0 (up sharply from Grok 4.6’s 20.3%); Artificial Analysis measures the same model at 26%. That 12-point harness gap is the single most important number in this comparison — if your workload is terminal-driven coding agents, use the independent figure, and compare it against Claude Sonnet 5.5’s 64% on AA before committing.

Its headline intelligence score is mid-pack. AA Intelligence Index v4.3.2 puts Grok 4.7 at 46, behind Fable 5.1 and GPT-6-class models.

The cost trap is tokens, not price. Per-token pricing is unchanged from Grok 4.6, but at xhigh effort 4.7 burns roughly 81,000 output tokens per task versus ~36,000 for 4.6. Same price list, materially higher bill. If you move from 4.6 to 4.7, re-measure cost per completed task rather than assuming parity.

Why Grok 4.8’s C++ story is the interesting part

Essentially every frontier lab trains on a Python-fronted stack over CUDA kernels — PyTorch or JAX. xAI claiming a from-scratch C++ training stack at 2.5 trillion parameters is a genuine divergence, and the bet is throughput: removing framework overhead and controlling the whole pipeline should buy you more useful FLOPs per dollar of hardware.

The counter-argument is the track record. Grok 4.7 reportedly missed at least five projected launches before shipping on September 21, 2026, and xAI then skipped ahead to naming 4.8 while 4.7 was still unreleased. From-scratch distributed-training infrastructure is exactly where silent numerical divergence, checkpoint corruption and gradient-sync bugs live — problems that cost months and do not surface in a blog post. Musk said on September 13–14 that pretraining would finish that week and RL would start immediately. As of October 3, 2026, three weeks later, there is still no model.

What to do

  • Shipping something this quarter: Grok 4.7, at high effort rather than xhigh unless you have measured that xhigh earns its 2× token burn. Watch the 200K-token pricing cliff.
  • Terminal-heavy coding agents: compare against Claude Sonnet 5.5 at the same $2 input price before choosing Grok — the independent Terminal-Bench gap is large.
  • Legal, EE, or multi-hour office automation: Grok 4.7 is genuinely competitive here and underrated for it.
  • Planning around Grok 4.8 or Grok 5: don’t. Neither has a date, and the 4.7 history says xAI’s dates move. Keep your model layer swappable and re-evaluate when a model page with a price exists.

Related: Claude Sonnet 5.5 vs GPT-6.1 Sol, Claude Opus 5.5 vs Grok 4.7 vs MiMo V2.6 Pro, best AI coding agent 2026.

Last verified: October 3, 2026. Prices from xAI docs; scores from Artificial Analysis and xAI’s published benchmarks.

Sources