OpenAI Ultrafast vs Fast Mode vs Claude Fast vs Grok Fast
The short answer
As of September 30, 2026, every frontier vendor sells speed as a separate line item: OpenAI Fast mode at 2x, OpenAI Ultrafast at 6x (up to 300 tokens/s, GPT-6 Astra $60/$300), Anthropic fast mode at 2x on Opus 5.5 ($8/$40, up to 2.5x speed), and xAI’s Grok 4.7 Fast at 2x ($4/$12, 2x speed). The cheapest fast frontier tokens today are Claude Sonnet 5.5 at ~138 tokens/s for the standard $2/$10 — no surcharge. Pay a speed tier only when a person is waiting on the output.
Speed tiers, priced
| Tier | Model | Input / output (per MTok) | Multiple | Claimed speed | Where |
|---|---|---|---|---|---|
| Standard | GPT-6 Astra | $10 / $50 | 1x | ~57 tok/s (AA) | API, ChatGPT |
| Fast mode | GPT-6 Astra | $20 / $100 | 2x | ”Faster than Standard” | API |
| Ultrafast | GPT-6 Astra | $60 / $300 | 6x | Up to 300 tok/s; 8x in Codex, 6x in API | API; Pro 500, Enterprise |
| Standard | GPT-6.1 Sol | $2 / $10 | 1x | ~67 tok/s (AA) | API, ChatGPT Work, Codex |
| Fast mode | GPT-6 Sol | $4 / $20 | 2x | — | API |
| Ultrafast | GPT-6.1 Sol | ~$12 / $60 (derived) | 6x | ”Coming days” | — |
| Standard | Claude Opus 5.5 | $4 / $20 | 1x | — | API, clouds |
| Fast mode | Claude Opus 5.5 | $8 / $40 | 2x | Up to 2.5x | API only |
| Standard | Claude Sonnet 5.5 | $2 / $10 | 1x | ~138 tok/s (AA) | API, clouds |
| Standard | Grok 4.7 | $2 / $6 | 1x | — | xAI API |
| Fast | Grok 4.7 Fast | $4 / $12 | 2x | 2x | xAI API |
| Standard | Gemini 3.8 Flash | $0.75 / $3.75 (intro) | 1x | — | Gemini API |
| Priority | Gemini 3.8 Flash | ~$1.35 / $6.75 | ~1.8x | — | Gemini API |
Ultrafast per-token rates for Sol are derived from OpenAI’s “6x standard” statement; OpenAI has not published a separate Sol Ultrafast line as of September 30, 2026. Sonnet 5.5 and GPT-6.x speeds are Artificial Analysis measurements at max effort.
The math on the reference task
30K input, 5K output, no cache:
| Cost | vs cheapest | |
|---|---|---|
| GPT-6.1 Sol standard | $0.11 | 1x |
| Sonnet 5.5 standard (~138 tok/s) | $0.11 | 1x |
| Grok 4.7 Fast | $0.18 | 1.6x |
| GPT-6 Sol Fast mode | $0.22 | 2x |
| Opus 5.5 standard | $0.22 | 2x |
| Opus 5.5 fast mode | $0.44 | 4x |
| GPT-6 Astra standard | $0.55 | 5x |
| GPT-6.1 Sol Ultrafast (derived) | $0.66 | 6x |
| GPT-6 Astra Fast mode | $1.10 | 10x |
| GPT-6 Astra Ultrafast | $3.30 | 30x |
The 30x spread between Sol standard and Astra Ultrafast is the whole story. A 5,000-token answer at 300 tokens/s takes ~17 seconds; at 57 tokens/s it takes ~88 seconds. Whether 70 seconds is worth $2.75 depends entirely on who is waiting.
How each vendor sells speed
OpenAI has three lanes. Fast mode (the renamed Priority processing) is a 2x queue-jump. Ultrafast is new silicon-backed throughput — OpenAI previewed it on GPT-5.6 Sol in August 2026 (then quoted at up to 750 tokens/s on Cerebras) and shipped the GA version at DevDay for GPT-6 Astra at “up to 300 tokens per second.” On stage: “300 tokens per second. You know what, it’s worth it.” The demo raced two agents building a rocket; the Ultrafast one launched first. In ChatGPT, only Pro 500 and Enterprise can select it, and credits on cheaper Pro tiers do not unlock it.
Anthropic sells fast mode only on Opus — Opus 5.5 at $8/$40 (API only, up to 2.5x), previously Opus 4.8 at $10/$50. Sonnet 5.5 gets no fast lane, but at ~138 tokens/s standard it is already faster than any OpenAI standard lane. Caveat from Artificial Analysis: Sonnet 5.5 at max effort emits ~193K output tokens per index task, so per-task wall clock can still be long — dial effort down for latency.
xAI sells Grok 4.7 Fast as a straight 2x/2x: double the price, double the speed, same 500K context. Since Grok 4.7 already emits ~81K output tokens per AA task, the Fast variant’s benefit is real for interactive use.
Google has no branded “ultrafast”; Priority tier is ~1.8x on Gemini 3.8 Flash, and Flash-class models are already fast (Gemini 3.5 Flash ~201 tok/s on AA). Google’s intro pricing on 3.8 Flash doubles to $1.50/$7.50 on January 1, 2027.
Specialists — diffusion and speculative models such as Mercury 2 (~769 tok/s) and Celeris-1 (~1,491 tok/s) on Artificial Analysis — beat every frontier lane on raw throughput but are not frontier-class on reasoning. For structured, low-stakes generation they are the cheapest way to be fast.
Decision rule
- Human typing at the model (Codex pairing, chat): Sonnet 5.5 standard first; if you need Astra-class reasoning at speed, Astra Fast mode before Ultrafast — 2x buys most of the felt improvement.
- Voice or live support agents: Ultrafast or a Flash/specialist model; latency is the product.
- Autonomous agents, batch, evals: never. Use GPT-6.1 Sol or Sonnet 5.5 with caching, Batch at 50% off, and parallelism. See cheapest AI API 2026 and prompt caching explained.
- Cheapest fast frontier tokens: Sonnet 5.5, no surcharge.
August’s view of the same race: GPT-5.6 Sol Ultrafast vs Gemini 3.7 Flash vs Grok 4.6.
Last verified: September 30, 2026. Vendor rate cards; Artificial Analysis speed measurements.