AI agents · OpenClaw · self-hosting · automation

Quick Answer

Together Link vs FireRouter vs OpenRouter for Claude Code

Published:

The short answer

All three let you keep Claude Code (and, for two of them, Codex) as your agent while changing what serves the tokens. They solve different problems:

Fireworks FireRouterTogether LinkOpenRouter
GoalCut Claude spend, keep Opus for hard turnsRun your agents on open modelsKeep Claude, add failover and budgets
RoutingPer user turn, cache-aware, between Claude/GPT and open modelsAuto router; Claude Opus for hard requests only in Claude sessions with an Anthropic keyYou choose the model; failover across Anthropic providers
HarnessesClaude Code, Codex, Cursor IDE and more (FireConnect CLI)Claude Code, Codex CLI, OpenCode 2, Pi Code 0.80.8+, Claude Desktop/Cowork, ChatGPT DesktopClaude Code via ANTHROPIC_BASE_URL
Open modelsGLM 5.3, GLM 5.3 Flash, Kimi K3, DeepSeek V4.1 FlashKimi K3, GLM 5.3, DeepSeek V4.1 Flash, Qwen 3.8Whole catalog, but Claude Code is only guaranteed on Anthropic first-party
BillingOpen turns at Fireworks serverless rates; closed turns on your own Anthropic/OpenAI keyTogether serverless rates; Opus on your Anthropic keyPass-through prices; 5.5% fee on credit purchases
StatusGA as a router model since Sep 28, 2026Beta, launched Oct 5, 2026GA
Published savings−57% cost/session at 98.1% of Opus accuracy (vendor A/B)“Over 50%” (vendor claim)None; cost control via budgets

Fireworks FireRouter

FireRouter gives you one model ID — firerouter/opus, bare firerouter, or a custom list such as firerouter/astra/opus/glm-5p3 — and picks a model for each new user turn. The chosen model serves the whole turn, including its tool calls. Today the Opus mix routes between Claude Opus 5.5, GLM 5.3 and GLM 5.3 Flash. It weighs the cost of losing the prompt cache before switching, which is why Fireworks calls it cache-aware.

Fireworks’ A/B test (published September 28, 2026) on its own coding traffic: cost per session fell from $15.36 to $6.63, graded accuracy 78.7% against 80.2% for Opus alone, cache hit rate 94.2% against 97.8%. A header (x-routing-preference, 1 = max intelligence to 5 = max savings, default 3) tunes the trade-off. Setup is two commands: fireconnect login then fireconnect claude --model firerouter/opus.

Limits: Claude and GPT turns run on your own Anthropic or OpenAI credential; FireRouter is not available through Microsoft Foundry or on accounts with data residency enabled. For open-models-only routing use auto (GLM 5.3, GLM 5.3 Flash, Kimi K3) or auto-instant.

One install (curl -fsSL https://link.together.ai/install | bash), then togetherlink claude, tcodex, topencode or tpi. Your normal configuration is untouched and going back takes one command. Prices per million tokens (input / cached / output): Kimi K3 $2.70 / $0.27 / $13.50, GLM 5.3 $1.40 / $0.26 / $4.40, DeepSeek V4.1 Flash $0.30 / $0.01 / $1.20, Qwen 3.8 $2 / $0.25 / $6. These are Together’s serving rates (the Kimi K3 figure is a promotion valid until October 11, 2026; standard $3/$15), not the model makers’ own APIs — DeepSeek’s first-party V4.1 Flash price is $0.15/$0.60 off-peak. In Claude Code, /model maps Opus → Kimi K3, Fable → GLM 5.3, Sonnet → DeepSeek V4.1 Flash.

The Auto router uses Claude Opus only in Claude Code and Claude Desktop sessions where you supplied an Anthropic key; Codex, OpenCode, Pi Code and ChatGPT Desktop always stay on Together models. A per-session tracker shows your spend next to what Opus 5.5 would have cost. It is a beta: commands and the model list may change.

OpenRouter

OpenRouter is not a cost router for Claude Code; it is a reliability and management layer. Point Claude Code at https://openrouter.ai/api with your key as ANTHROPIC_AUTH_TOKEN and requests fail over between Anthropic providers when one is down or rate-limited, with team spending limits and a usage dashboard. Inference has no markup; credit purchases carry 5.5% ($0.80 minimum). BYOK is free up to $25,000 of list-price usage a month. OpenRouter states Claude Code is only guaranteed to work with the Anthropic first-party provider, so do not expect clean results from non-Claude models through it.

Which to pick

  1. Claude Code is your agent and the bill is the problem: FireRouter with Opus. It is the only one with a published quality-versus-cost measurement.
  2. Your team uses several harnesses and you want open models everywhere: Together Link.
  3. You must stay on Claude (policy, quality) but need uptime and budgets: OpenRouter.
  4. You need data residency: none of the routers above — pin a residency-compatible model directly.

Model list prices are in our current API prices table (Claude Opus 5.5: $4/$20 per million tokens). For choosing the provider behind the open models, see best LLM serving providers for coding assistants.

Last verified: October 9, 2026.

Sources