AI agents · OpenClaw · self-hosting · automation

Quick Answer

What Is Grok 4.7? xAI's New Model Explained (Sep 2026)

Published:

The short answer

Grok 4.7 is xAI’s new coding-and-knowledge-work model, released September 21, 2026, at exactly Grok 4.6’s price: $2 per million input tokens and $6 per million output tokens. xAI calls it “our most capable model for coding and knowledge work” and, more carefully, “highly competitive in its class.” Independent numbers agree with the careful version: it is a large step up from Grok 4.6 and the cheapest model in its tier, but it sits clearly behind Claude Fable 5.1 and GPT-6 Astra on agentic coding.

What xAI shipped on September 21, 2026

ItemGrok 4.7 (per xAI’s launch post and model docs)
ReleaseSeptember 21, 2026 (x.ai/news post, docs entry and @SpaceXAI post within the same hour)
Base model”A new, larger base model” than Grok 4.6; no parameter count published
TrainingLonger reinforcement-learning run on a harder task mix “weighted toward problems that take many hours to complete”
HarnessTrained to natively understand the Grok Bot harness
Context window500,000 tokens
Price, under 200K prompt tokens$2.00 input / $0.50 cached / $6.00 output per million
Price, 200K and above$4.00 input / $1.00 cached / $12.00 output per million
Fast variantTwice the output speed at twice the price
Knowledge cutoffMay 2026
AvailabilityCursor, Grok Build (free to try), xAI API, GitHub Copilot (rolling out Sept 21), Vercel AI Gateway (40% off for the launch week), model routers
Safety62.4% on LatchBio’s biosafety benchmark; 3.3% of risky dual-use prompts allowed through on xAI’s HackerBench v0.3; invite-only red-team access for cybersecurity partners

Elon Musk had claimed on July 28, 2026 that Grok 4.7 would be a 2.1-trillion-parameter model trained partly on SpaceX engineering data. Neither claim appears in xAI’s launch material. The launch itself came 31 days after the first “in 4 weeks” target Musk posted on July 24.

xAI’s benchmark table

xAI published one table comparing Grok 4.7 at xHigh effort with Grok 4.6 High, GPT-5.6 Sol Max and Claude Fable 5.1 Max. These are xAI-run numbers.

BenchmarkGrok 4.7Grok 4.6GPT-5.6 SolFable 5.1
Input / output price per MTok$2 / $6$2 / $6$4 / $20$10 / $50
CursorBench 4.0 (software engineering)46.3%40.4%41.7%51.8%
DeepSWE v1.1 (software engineering)71.0% (high)65.2%72.7%70.0%
EEBench (electrical engineering)64.0%53.0%39.4%56.4%
AA Briefcase v1.1 (multi-hour office work)1,6571,5461,4871,678
Terminal-Bench 4.0 (multi-hour terminal work)38.0%20.3%37.3%57.9%
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional (clinical reasoning)56.7%48.5%60.5%62.1%

Read across: Grok 4.7 beats Grok 4.6 on all seven rows, with Terminal-Bench nearly doubling. Against GPT-5.6 Sol it leads on four of seven. Against Fable 5.1 Max it leads on two (EEBench, Harvey Legal), sits within a point on two (DeepSWE, AA Briefcase) and trails on the two coding rows Cursor users care about most. On xAI’s GDPval chart for professional knowledge work, Fable 5.1 scored 1,735 Elo, Grok 4.7 1,695, Grok 4.6 1,605 and GPT-6 Astra 1,542.

The first independent reads

Artificial Analysis posted its result 21 minutes after the announcement:

  • Intelligence Index v4.3.2: 46 for Grok 4.7 (xhigh), rank 16 of 655 models. Claude Fable 5.1 and GPT-6 both score 53. Grok 4.7’s two highest reasoning levels score about the same.
  • Terminal-Bench 4.0: 26%, versus 60% for GPT-6 Astra, 55% for Claude Fable 5.1 and 27% for the far cheaper DeepSeek V4.1 Flash. This is the widest gap in the release.
  • Token usage: ~81,000 output tokens per index task, against ~36,000 for Grok 4.6 at the same effort. Artificial Analysis’s framing: “Grok 4.7’s gains come with higher token usage.”
  • Speed: ~188 tokens per second on long prompts, about 7.1 minutes per index task.
  • AA-Omniscience hallucination rate: 29%, down from 34% for Grok 4.6 (high); accuracy 47% versus 48%.

Cursor confirmed availability at “the same price and speed as Grok 4.6” and kept Grok 4.6 on its model menu.

What the price actually buys

The $2/$6 rate card is the story. At list price Grok 4.7 is 2x cheaper than GPT-5.6 Sol and 5x cheaper than Fable 5.1 or GPT-6 Astra per token. But because it thinks roughly twice as long as Grok 4.6 to earn its scores, a task that cost one unit on 4.6 costs about two on 4.7 — still far below the frontier pair per task, but not the 5x headline suggests once token counts are included. OpenRouter listed it at $1.60/$4.80 on launch day, about 20% under xAI direct.

Who should use Grok 4.7

  • Yes: Cursor and Grok Build users on cost-sensitive agentic coding; electrical-engineering and legal-workflow tasks where its two benchmark leads are relevant; anyone already on Grok 4.6 (drop-in, same price; see Grok 4.7 vs Grok 4.6: should you upgrade).
  • No: long multi-hour terminal agents where Terminal-Bench 4.0 predicts outcomes — GPT-6 Astra and Claude Fable 5.1 are 2x better there; clinical reasoning; anything where a 29% hallucination rate on AA-Omniscience is disqualifying.

Full three-way breakdown: Grok 4.7 vs Claude Fable 5.1 vs GPT-6 Astra vs GPT-5.6 Sol.

Sources