AI agents · OpenClaw · self-hosting · automation

Quick Answer

Cheapest Inference Provider for Open-Weight Models 2026

Published:

The short answer

For most open-weight models in October 2026, DeepInfra is the cheapest mainstream inference provider per million tokens, often at a third to a quarter of the model maker’s list price. Together AI, Fireworks and Baseten typically match the maker’s list price and sell reliability, throughput and dedicated capacity instead. Groq and Cerebras charge more and sell speed. The model maker’s own API is competitive only in specific cases, chiefly DeepSeek off-peak. Prices below were read from the OpenRouter provider endpoints API and vendor pages on October 5, 2026; check precision (BF16, FP8, FP4) before comparing.

Price per million tokens (input / output), October 5, 2026

ModelCheapest mainstreamModel makerTogetherFireworksGroq / Cerebras
gpt-oss-120b (OpenAI, open weights)DeepInfra $0.037 / $0.17 (BF16)—$0.15 / $0.60—Groq $0.15 / $0.60 · Cerebras $0.35 / $0.75
DeepSeek V4.1 FlashDeepInfra $0.14 / $0.42 (FP8)DeepSeek $0.15 / $0.60 off-peak, $0.30 / $1.20 peak$0.30 / $1.20$0.30 / $1.20—
GLM-5.3 (Z.ai)Novita $0.42 / $1.32 (FP8) · DeepInfra $0.562 / $2.50 (FP4)Z.ai $1.40 / $4.40$1.40 / $4.40$1.40 / $4.40—
Kimi K3 (Moonshot)DeepInfra $2.85 / $14.25 (MXFP4)Moonshot $3 / $15$2.70 / $13.50$3 / $15—

Two patterns stand out. First, the discount is largest on smaller, older models such as gpt-oss-120b, where serving cost has fallen furthest. Second, on new frontier-scale open models such as Kimi K3 almost everyone charges close to the maker’s price, because few hosts have enough capacity to undercut it.

Provider by provider

DeepInfra — cheapest default. Consistently at or near the bottom of the price table, with an OpenAI-compatible API and a broad catalogue. Check the precision label: its GLM-5.3 endpoint is FP4, its gpt-oss-120b is BF16. Good for batch jobs, prototyping and cost-sensitive production.

DeepSeek (first-party) — cheapest off-peak for its own models. V4.1 Flash at $0.15 / $0.60 off-peak is close to DeepInfra on input and includes a $0.003 cache-hit price that no reseller matches. Peak hours (01:00–04:00 and 06:00–10:00 UTC) double the price, which hurts Asia-Pacific daytime workloads.

Together AI, Fireworks, Baseten — maker’s price, better operations. They charge list price on most models and win on rate limits, fine-tuning, dedicated GPUs and enterprise contracts. Our Baseten vs Together vs Fireworks comparison covers the operational differences.

Groq and Cerebras — pay more for speed. Both cost more per token than the cheapest hosts on gpt-oss-120b; choose them for latency-critical chat and voice agents.

OpenRouter — one key, many providers. OpenRouter routes to these same endpoints and lets you sort by price or pin a provider; useful for comparing live prices and for failover. See OpenRouter vs Together vs Fireworks.

Read the fine print before you switch

  • Precision. FP4 and FP8 endpoints are cheaper and can be measurably worse on long reasoning or code. Evaluate the exact endpoint.
  • Output price. Agents and reasoning models emit far more output than input; some endpoints with near-zero input prices charge $2 or more for output.
  • Context length. A cheap endpoint with a 128K or 256K window will fail on prompts that the 1M-token listing accepts.
  • Caching. Prompt-cache discounts (DeepSeek’s $0.003 cache-hit rate, for example) can matter more than the headline price for agent loops.

The pick: DeepInfra for the lowest cost per token; DeepSeek’s own API off-peak for DeepSeek models; Together or Fireworks when you need guaranteed throughput; Groq when you need speed. For the full provider ranking beyond price, see best LLM inference providers 2026, and our current API prices table for first-party model rates.

Last verified: October 5, 2026. Third-party prices from the OpenRouter provider endpoints listing and DeepInfra’s pricing page; first-party prices from DeepSeek’s API documentation. Prices change often; recheck before committing.

Sources