AI agents · OpenClaw · self-hosting · automation

Quick Answer

DeepSeek V4 Flash 0731 vs Kimi K3 vs GLM-5.2: Best Open-Weight Coding (August 2026)

Published:

The Short Answer

For open-weight coding in August 2026: DeepSeek V4 Flash 0731 is the cheapest via API and strong on agents; Kimi K3 (weights opened July 27, 2026) and GLM-5.2 are the picks when you need to run the weights yourself. All three are Chinese-lab open-weight families closing on the frontier.

The Comparison

DeepSeek V4 Flash 0731Kimi K3GLM-5.2
API price (per MTok)$0.14 / $0.28*$3 / $15 (flat)Low single digits
WeightsReleasedOpened Jul 27, 2026Downloadable
Context1MLargeLarge
Terminal-Bench 2.182.7CompetitiveCompetitive
Best atCheapest hosted APISelf-host frontier-adjacentSelf-host general coding

*DeepSeek doubles during peak hours.

Where Each Wins

  • DeepSeek V4 Flash 0731 → cheapest managed API. Native Codex + Responses API support, DSpark speculative decoding, 82.7 Terminal-Bench 2.1. Almost always the cheapest way to run a capable coding agent.
  • Kimi K3 → self-host frontier-adjacent. Weights opened July 27, 2026; flat $3/$15 if you use the hosted API. Strong choice when you want control and no peak/off-peak pricing games.
  • GLM-5.2 → self-host general coding. Downloadable weights and solid all-round coding make it a dependable on-prem option.

Hosted API vs Self-Host

  • Cheapest bill, no ops → DeepSeek V4 Flash 0731 hosted API
  • Data control / on-prem → Kimi K3 or GLM-5.2 weights on your own GPUs
  • Predictable flat pricing → Kimi K3

Self-hosting trades per-token fees for GPU cost and ops overhead — worth it at high volume or under strict data-residency rules.

Verdict

  • Cheapest capable open-weight API → DeepSeek V4 Flash 0731
  • Best self-host, flat pricing → Kimi K3
  • Best self-host general coding → GLM-5.2

Sources