AI agents · OpenClaw · self-hosting · automation

Quick Answer

Kimi K3 vs Qwen 3.7 Max vs DeepSeek V4 Pro (2026)

Published:

Kimi K3 vs Qwen 3.7 Max vs DeepSeek V4 Pro (2026)

The open-weight (and near-open) frontier has three standouts in July 2026: Kimi K3 (peak capability), Qwen 3.7 Max (coding-agent value), and DeepSeek V4 Pro (cheapest at scale). Here’s how they compare and when each wins.

Last verified: July 25, 2026

Head to Head

Kimi K3Qwen 3.7 MaxDeepSeek V4 Pro
VendorMoonshot AIAlibabaDeepSeek
LicenseOpen weights (Jul 27)Proprietary APIMIT (open weights)
Params2.8T MoEUndisclosed (frontier)1.6T total / 49B active
Intelligence Index~57 (#3 overall)~56.644
Input / Output (per MTok)$3 / $15 (flat)$1.25 / $3.75 (promo)$0.435 / $0.87 (off-peak)
Cost / task (30K→5K)~$0.165~$0.056~$0.017
Context1M tokens1M tokensLarge
Best forPeak capability + visionCoding agentsHigh-volume text at low cost

Kimi K3 — peak open capability

Moonshot AI announced Kimi K3 — a 2.8-trillion-parameter MoE with a 1M-token context — on July 16, 2026, with open weights scheduled for July 27, 2026. On the Artificial Analysis Intelligence Index it scores about 57, ranking #3 overall behind only Claude Fable 5 and GPT-5.6 Sol, and comparable to Opus 4.8 and GPT-5.5. It also brings native vision. Pricing is a flat $3/$15 per MTok (no peak/valley). Pick K3 when you want the strongest open model, difficult judgment, or multimodal work.

Qwen 3.7 Max — coding-agent value

Alibaba’s Qwen 3.7 Max is a proprietary frontier agent model (1M context) tuned for long-horizon agentic coding. It posts SWE-Verified ~80.4, Terminal-Bench 2.0 69.7, and SWE-Pro 60.6 — on par with or ahead of Opus 4.6 — and demonstrated 35+ hours of continuous autonomous coding with 1,000+ tool calls. It’s natively compatible with both OpenAI and Anthropic API specs, so it drops into Claude Code. At $1.25/$3.75 on its 50% promo (list $2.50/$7.50), plus a 90%-off cached input tier, it’s excellent value. Pick Qwen 3.7 Max when you’re running coding agents and want frontier coding quality cheaply.

DeepSeek V4 Pro — cheapest at scale

DeepSeek V4 Pro (1.6T total / 49B active, MIT-licensed) leads open weights for raw cost efficiency. It scores 44 on the Intelligence Index — lower than K3 — but its off-peak pricing of $0.435/$0.87 (doubling during peak UTC hours) makes a real task cost about $0.017, roughly a tenth of K3. Cache hits drop to ~$0.043/MTok. Pick V4 Pro when cost, volume, and predictable text workflows lead the decision.

The Winning Pattern: Route by Job

  • Default to DeepSeek V4 Pro for high-volume, cost-sensitive text.
  • Escalate to Qwen 3.7 Max for coding agents and tool-heavy work.
  • Escalate to Kimi K3 for the hardest reasoning, vision, or judgment.

Because output tokens dominate cost, a router that reserves K3 for the tasks that need it typically cuts spend sharply versus running one model on everything.

Bottom Line

  • Cheapest per task: DeepSeek V4 Pro (~$0.017, MIT-licensed)
  • Best coding-agent value: Qwen 3.7 Max (~$0.056)
  • Peak open capability: Kimi K3 (~$0.165, #3 overall, open weights Jul 27)

Benchmark on your real prompts — the open frontier is now close enough to US flagships that many teams can run these models as primary, not fallback.

Sources