AI agents · OpenClaw · self-hosting · automation

Quick Answer

Kimi K3 vs GLM-5.2 vs DeepSeek V4 Pro: Cost (July 2026)

Published:

The Short Answer

Among open trillion-scale MoE models in July 2026, DeepSeek V4 Pro is the cheapest by a wide margin, GLM-5.2 is the mid-priced middle ground, and Kimi K3 is the premium option — the first open 3T-class model with native vision, but ~17x DeepSeek’s output cost.

Head-to-Head

Kimi K3GLM-5.2DeepSeek V4 Pro
MakerMoonshot AIZhipuDeepSeek
Params (total)~2.8T MoE744B1.6T (49B active)
API in/out ($/MTok)$3 / $15~ / ~$4.40 out$0.43 / $0.87
Context1Mlarge1M
MultimodalYes (native vision)Text-focusedText-only
LicenseModified MITMITMIT-style
Blended cost/task~$0.94mid~$0.04

Where Each Wins

  • Cheapest to serve → DeepSeek V4 Pro. ~$0.43/$0.87 per MTok and roughly $0.04 per task on blended cost — an order of magnitude below Kimi K3’s ~$0.94. Text-only, 1M context, permissive license. The default open pick for cost-sensitive production.
  • Best value middle → GLM-5.2. 744B, MIT-licensed, ~$4.40 output. Cheaper to serve than K3 while staying capable — a sensible pick if DeepSeek feels too minimal but K3 is overkill.
  • Top open intelligence + multimodal → Kimi K3. The first open 3T-class model, with native vision, always-on reasoning, and a 1M context. It edges DeepSeek on some intelligence indexes — but at $3/$15 it’s a premium-tier open model, ~17x DeepSeek V4 Pro’s output cost.

The Cost Gap Is Huge

On output tokens alone, Kimi K3 ($15) is roughly 17x DeepSeek V4 Pro ($0.87). Unless you specifically need vision or K3’s peak reasoning, that gap is very hard to justify at scale. GLM-5.2 sits in between and is the pragmatic compromise for teams that want more than DeepSeek without K3’s bill.

How To Decide

  • Pure cost-per-task, text-only → DeepSeek V4 Pro.
  • Need vision or top open intelligence → Kimi K3.
  • Want a capable, cheaper-than-K3 middle → GLM-5.2.
  • Self-hosting? All three are open-weight; K3’s 3T scale needs the most hardware, DeepSeek V4 Pro the least among these for a given throughput.

Sources


Last verified: July 30, 2026