Fugu Max vs Fugu Ultra v2 vs Claude Sonnet 5: Which to Pick
The short answer
| Fugu Max v1.0 | Fugu Ultra v2.0 | Claude Sonnet 5 | |
|---|---|---|---|
| Maker | Sakana AI (Tokyo) | Sakana AI | Anthropic |
| Released | September 11, 2026 | September 11, 2026 (v2.0) | GA June 30, 2026 |
| What it is | Learned orchestrator over Sakana’s largest model pool, tuned for cost-performance | Learned orchestrator over a deeper expert pool, tuned for quality | Single frontier model |
| Input / output (per MTok) | $2 / $6, flat at any context | $5 / $30 (≤272K) · $10 / $45 (>272K) | $2 / $10 |
| Cached input | $0.25 | $0.50 (≤272K) · $1.00 (>272K) | $0.20 |
| Extras | web_search / web_fetch $0.007 per call | — | Tools billed as tokens |
| Benchmark claim | Best overall on Terminal-Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench, SWEFish | Best or joint-best on 5 of 8 (GDP.pdf, Chartography, DeepSWE, Toolathon, SWEFish); top-2 on 7 of 8 | Anthropic’s mid-tier workhorse; strong on coding, 1M context |
| Frontier models in pool | Not Fable 5 / 5.1 or GPT-6 Astra | Same exclusion; cutoff Aug 28, 2026 | n/a |
| EU/EEA availability | No | No | Yes |
| Best for | High-volume coding, review, research where output tokens dominate cost | Kaggle-grade problems, paper reproduction, security analysis, patent research | Predictable single-vendor behaviour, EU users, Anthropic ecosystem |
Pick Fugu Max if output tokens dominate your bill and you are outside the EU; pick Fugu Ultra v2 when a hard problem justifies 5x the output price; pick Claude Sonnet 5 when you need one accountable model, EU availability, or Anthropic-native tooling.
What Sakana shipped on September 11, 2026
Sakana AI released Fugu Max v1.0 and Fugu Ultra v2.0 on September 11, 2026, expanding the Fugu line to four models — Fugu, Fugu Ultra, Fugu Max and Fugu Cyber — all behind one OpenAI-compatible API. Fugu is not a single network. It is a learned coordinator (grounded in Sakana’s two ICLR 2026 papers, TRINITY and the Conductor) that assembles Thinker/Worker/Verifier roles from a pool of open and specialised models and routes each step of a task to whichever agent it predicts will do it best. Sakana’s pitch is “frontier-level performance without single-vendor dependency.”
Pricing is deliberately simple. Sakana says it “never stacks model fees”: when multiple agents are active you pay a single rate based on the top-tier model involved, and for Fugu Max that rate is a flat $2 input / $6 output per million tokens regardless of context length. Fugu Ultra v2.0 is $5 / $30 up to 272K tokens and $10 / $45 above it, with cached input at $0.50 / $1.00.
Fugu Max: the cost-performance play
Fugu Max orchestrates Sakana’s largest pool to date, integrating “an unprecedented number of open-weights and specialized models” including NVIDIA’s Nemotron family via a Sakana–NVIDIA collaboration. Sakana’s claims for it:
- Best overall score on six benchmarks: Terminal-Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench and SWEFish (Sakana’s internal coding benchmark).
- Output pricing 40–60% lower than Claude Sonnet 5, GPT-5.6 Terra and Kimi K3.
- Expands the cost-performance Pareto frontier on seven of ten benchmarks.
Do the arithmetic on a code-review workload of 1M input and 1M output tokens per day: Fugu Max is $8/day, Sonnet 5 $12, GPT-5.6 Terra $14, Kimi K3 $18. The delta is entirely output price — inputs are identical to Sonnet 5 and Terra.
Fugu Ultra v2.0: the quality tier
Fugu Ultra coordinates a deeper pool of expert agents for “hard, high-stakes problems”; early users cite Kaggle competitions, paper reproduction, cybersecurity analysis and literature/patent investigations. Sakana claims v2.0 is best or joint-best on five of eight benchmarks (GDP.pdf, Chartography, DeepSWE, Toolathon, SWEFish) and top-2 on seven of eight. Its published demo is a mechanical CAD iris where Fugu Ultra’s blades actually close the aperture and competitors’ do not.
The most important disclosure is in the footnote: Fugu Ultra’s training cutoff is August 28, 2026, and Claude Fable 5, Fable 5.1 and GPT-6 Astra are not in its pool. Sakana frames this as a feature — Fugu “does not rely on individual proprietary frontier models to deliver frontier output” and therefore protects users from “vendor lock-in, API revocations, geopolitical turbulence, and sudden service cutoffs.” It also means Ultra’s ceiling is set by the open and specialised models it can reach, not by the newest closed flagships.
Claude Sonnet 5: the single-model baseline
Anthropic’s Claude Sonnet 5 went GA on June 30, 2026 at $2 / $10 per MTok, and Anthropic cancelled the scheduled September 1 increase to $3/$15 in August 2026, making $2/$10 the standard price. Cache hits cost $0.20, five-minute cache writes $2.50 and one-hour writes $4. What Sonnet 5 offers that Fugu does not:
- One model, one behaviour. No orchestration layer choosing a different sub-model on Tuesday than it did on Monday; easier evals, easier prompt caching, easier incident review.
- EU/EEA availability and Anthropic’s enterprise compliance surface (zero-data-retention options, Bedrock/Vertex/Foundry deployment).
- Anthropic-native tooling: Claude Code, the Agent SDK, MCP-first integrations.
What it gives up: a higher output price than Fugu Max and no ability to swap providers when Anthropic changes a price, deprecates a model or tightens a safeguard.
Where the orchestration model is risky
- Determinism and evals. An orchestrator that “dynamically assembles agents” is harder to regression-test than one model. Pin a model ID, snapshot your eval suite, and re-run it whenever Sakana announces a pool change.
- Data residency. Your prompt may be forwarded to any model in the pool. Sakana lets you opt providers out for compliance — use it, and ask which vendors remain.
- Benchmark provenance. All Fugu numbers are Sakana-run. SWEFish is internal. Independent leaderboards had not yet scored v2.0 or Max at publication.
- Geography. No EU/EEA service as of September 2026.
Pricing plans beyond the API
Sakana also sells subscriptions for hands-on use — Standard $20/month, Pro $100/month, Max $200/month — each including Fugu, Fugu Ultra and Fugu Max with usage allowances; Sakana recommends pay-as-you-go tokens for “serious workloads.” Fugu Cyber (86.9% on Sakana’s cited security benchmark) is sales-only. Fugu is also listed on OpenRouter and LLM Gateway as of September 11, 2026.