Together Link vs FireRouter vs OpenRouter for Claude Code
The short answer
All three let you keep Claude Code (and, for two of them, Codex) as your agent while changing what serves the tokens. They solve different problems:
| Fireworks FireRouter | Together Link | OpenRouter | |
|---|---|---|---|
| Goal | Cut Claude spend, keep Opus for hard turns | Run your agents on open models | Keep Claude, add failover and budgets |
| Routing | Per user turn, cache-aware, between Claude/GPT and open models | Auto router; Claude Opus for hard requests only in Claude sessions with an Anthropic key | You choose the model; failover across Anthropic providers |
| Harnesses | Claude Code, Codex, Cursor IDE and more (FireConnect CLI) | Claude Code, Codex CLI, OpenCode 2, Pi Code 0.80.8+, Claude Desktop/Cowork, ChatGPT Desktop | Claude Code via ANTHROPIC_BASE_URL |
| Open models | GLM 5.3, GLM 5.3 Flash, Kimi K3, DeepSeek V4.1 Flash | Kimi K3, GLM 5.3, DeepSeek V4.1 Flash, Qwen 3.8 | Whole catalog, but Claude Code is only guaranteed on Anthropic first-party |
| Billing | Open turns at Fireworks serverless rates; closed turns on your own Anthropic/OpenAI key | Together serverless rates; Opus on your Anthropic key | Pass-through prices; 5.5% fee on credit purchases |
| Status | GA as a router model since Sep 28, 2026 | Beta, launched Oct 5, 2026 | GA |
| Published savings | −57% cost/session at 98.1% of Opus accuracy (vendor A/B) | “Over 50%” (vendor claim) | None; cost control via budgets |
Fireworks FireRouter
FireRouter gives you one model ID — firerouter/opus, bare firerouter, or a custom list such as firerouter/astra/opus/glm-5p3 — and picks a model for each new user turn. The chosen model serves the whole turn, including its tool calls. Today the Opus mix routes between Claude Opus 5.5, GLM 5.3 and GLM 5.3 Flash. It weighs the cost of losing the prompt cache before switching, which is why Fireworks calls it cache-aware.
Fireworks’ A/B test (published September 28, 2026) on its own coding traffic: cost per session fell from $15.36 to $6.63, graded accuracy 78.7% against 80.2% for Opus alone, cache hit rate 94.2% against 97.8%. A header (x-routing-preference, 1 = max intelligence to 5 = max savings, default 3) tunes the trade-off. Setup is two commands: fireconnect login then fireconnect claude --model firerouter/opus.
Limits: Claude and GPT turns run on your own Anthropic or OpenAI credential; FireRouter is not available through Microsoft Foundry or on accounts with data residency enabled. For open-models-only routing use auto (GLM 5.3, GLM 5.3 Flash, Kimi K3) or auto-instant.
Together Link
One install (curl -fsSL https://link.together.ai/install | bash), then togetherlink claude, tcodex, topencode or tpi. Your normal configuration is untouched and going back takes one command. Prices per million tokens (input / cached / output): Kimi K3 $2.70 / $0.27 / $13.50, GLM 5.3 $1.40 / $0.26 / $4.40, DeepSeek V4.1 Flash $0.30 / $0.01 / $1.20, Qwen 3.8 $2 / $0.25 / $6. These are Together’s serving rates (the Kimi K3 figure is a promotion valid until October 11, 2026; standard $3/$15), not the model makers’ own APIs — DeepSeek’s first-party V4.1 Flash price is $0.15/$0.60 off-peak. In Claude Code, /model maps Opus → Kimi K3, Fable → GLM 5.3, Sonnet → DeepSeek V4.1 Flash.
The Auto router uses Claude Opus only in Claude Code and Claude Desktop sessions where you supplied an Anthropic key; Codex, OpenCode, Pi Code and ChatGPT Desktop always stay on Together models. A per-session tracker shows your spend next to what Opus 5.5 would have cost. It is a beta: commands and the model list may change.
OpenRouter
OpenRouter is not a cost router for Claude Code; it is a reliability and management layer. Point Claude Code at https://openrouter.ai/api with your key as ANTHROPIC_AUTH_TOKEN and requests fail over between Anthropic providers when one is down or rate-limited, with team spending limits and a usage dashboard. Inference has no markup; credit purchases carry 5.5% ($0.80 minimum). BYOK is free up to $25,000 of list-price usage a month. OpenRouter states Claude Code is only guaranteed to work with the Anthropic first-party provider, so do not expect clean results from non-Claude models through it.
Which to pick
- Claude Code is your agent and the bill is the problem: FireRouter with Opus. It is the only one with a published quality-versus-cost measurement.
- Your team uses several harnesses and you want open models everywhere: Together Link.
- You must stay on Claude (policy, quality) but need uptime and budgets: OpenRouter.
- You need data residency: none of the routers above — pin a residency-compatible model directly.
Model list prices are in our current API prices table (Claude Opus 5.5: $4/$20 per million tokens). For choosing the provider behind the open models, see best LLM serving providers for coding assistants.
Last verified: October 9, 2026.