AI agents · OpenClaw · self-hosting · automation

Quick Answer

Open-Weight vs Proprietary LLMs 2026: When to Switch

Published:

The short answer

Open-weight models won the volume war in August 2026 and are nowhere near winning the spend war, and both facts are correct at once. They ran 56% of tokens on Vercel’s AI Gateway (7% in December 2025) for 14% of spend, because their tokens cost about 7.8× less. Enterprises far beyond Silicon Valley are routing routine work to them — AT&T 40% of employee requests heading to 60-70%, Coinbase reporting ~50% savings — while keeping frontier closed models for the tasks where a 7-12 point intelligence gap costs more than the tokens save. The right question for 2026 is not “open or closed” but “what fraction of my calls need the top of the index.”

The data, as of September 2026

MetricValueSource
Open-weight share of Vercel AI Gateway tokens, Aug 202656% (36% Jul; 13% Apr; 7% Dec 2025)Vercel Production Index, Sep 17
Open-weight share of gateway spend14%Vercel
Implied cost ratio, closed : open tokens~7.8 : 1Techstrong analysis of Vercel data
Avg price per token, Aug vs Jul−23.2% (third straight monthly drop)Vercel
Anthropic share of gateway spend64% (≥61% every month since Dec)Vercel
AT&T requests on open models40%, target 60-70%; ~45B tokens/dayFT, Sep 27
AT&T inference cost reduction~56%FT
Coinbase inference cost reduction~50%FT
Open-model mentions on US earnings calls6× year over yearFT
Top open model, AA Intelligence IndexMiMo-V2.6-Pro 46 (Opus 5.5 58, Fable 5.1 / GPT-6 Astra 53)Artificial Analysis

Why volume moved: the price gap is real

Per-million-token list prices, September 2026:

ModelWeightsInputOutput
Naive-N0.5-FlashMIT$0.10$0.40
DeepSeek V4.1 FlashMIT$0.15 (off-peak)$0.60 (off-peak)
Qwen 3.8 FlashOpen$0.15$0.47
MiMo-V2.6-ProMIT$0.435$0.87
GLM-5.3Staged, not shipped$1.40$4.40
Kimi K3Open$3$15
GPT-6 LunaClosed$0.10$0.50
Gemini 3.8 FlashClosed$0.75 (intro)$3.75 (intro)
Claude Sonnet 5Closed$2$10
Claude Opus 5.5Closed$4$20
GPT-6 Astra / Claude Fable 5.1Closed$10$50

Note the exception: OpenAI’s GPT-6 Luna at $0.10/$0.50 is priced like an open model. The cheap tier is no longer exclusively open-weight; what open weights add is the option to self-host, negotiate with multiple inference vendors, and avoid a single provider’s repricing. Full table in GPT-6 Luna vs Gemini 3.8 Flash vs DeepSeek V4.1 Flash.

Why spend stayed closed: the quality gap is also real

On the current Artificial Analysis Intelligence Index, Claude Opus 5.5 leads at 58, Claude Fable 5.1 and GPT-6 Astra sit at 53, GPT-6 Sol at 48, and the best open-weight model — Xiaomi’s MiMo-V2.6-Pro, 1.02T MoE, MIT — is at 46, with DeepSeek V4.1 Flash at 39. Seven to twelve points is not much on a single classification call. It is decisive on a 40-step agentic coding task where each step compounds. That is why Anthropic kept 64% of gateway spend even as its own customers moved from Fable 5 to Opus 5 — they stepped down within the closed tier, not out of it. Vercel’s data shows nine in ten Fable teams cut usage in August and more moved to Opus 5 than to any other model.

Three catches before you switch

1. Open-weight API prices are not stable. DeepSeek raised API prices 2.3-4.5× in August 2026 — its CEO credited the hike for crossing $1B annualized revenue — and introduced peak/off-peak windows where the same call costs double from 01:00-04:00 and 06:00-10:00 UTC. Z.ai’s GLM-5.3 Flash launch promo halved the list price until September 9, then expired. Free weights do not mean fixed prices; they mean you can switch vendors when prices move.

2. Cost per task beats cost per token. Artificial Analysis measured Gemini 3.8 Flash consuming ~123M output tokens to run its full index versus ~16M for GPT-6 Astra — a 7.7× gap that erases most of a 13× price advantage. Cheap reasoning models talk more. Benchmark your own tasks end to end; see token efficiency vs token price for the single-GPU angle.

3. Harnesses still assume closed models. Vercel’s CEO noted on September 2026 data that enterprise adoption is “still early, and harnesses, CLIs, IDEs, SDKs etc. need to be adapted to be model-agnostic.” If your coding agent, eval suite and prompts were tuned on Claude, the first week on Qwen will look worse than the model is.

A routing rule for 2026

Route by task class, and let the fraction drift toward open as models improve:

  • Open-weight by default: classification, tagging, extraction with a null-safe schema, summarization, translation, embeddings, first-draft code, test generation, internal chat over documents. This is the 56%.
  • Closed frontier by default: multi-hour agentic coding, anything where one wrong step costs an engineer-day, legal and financial drafting that ships without review, tasks where you have measured the gap and it matters.
  • Cheap closed (GPT-6 Luna, Gemini 3.8 Flash) as a middle tier when you want low price without running your own inference.
  • Self-host only when you have a compliance reason, a stable load above roughly 1B tokens a day, or an engineer who owns the serving stack. Below that, open-weight models via a gateway get you the price without the operations.
  • Re-evaluate quarterly. The open-weight share on Vercel rose every month from April to August 2026. The line will keep moving.

For the open flash tier specifically, see Naive-N0.5-Flash vs MiMo-V2.6-Flash vs DeepSeek V4.1 Flash vs GLM-5.3 Flash.

Last verified: September 28, 2026. Gateway data is Vercel AI Gateway through August 2026; prices from vendor pages.

Sources