Quick Answer
DeepSeek V4 Flash 0731 vs Kimi K3 vs GLM-5.2: Best Open-Weight Coding (August 2026)
The Short Answer
For open-weight coding in August 2026: DeepSeek V4 Flash 0731 is the cheapest via API and strong on agents; Kimi K3 (weights opened July 27, 2026) and GLM-5.2 are the picks when you need to run the weights yourself. All three are Chinese-lab open-weight families closing on the frontier.
The Comparison
| DeepSeek V4 Flash 0731 | Kimi K3 | GLM-5.2 | |
|---|---|---|---|
| API price (per MTok) | $0.14 / $0.28* | $3 / $15 (flat) | Low single digits |
| Weights | Released | Opened Jul 27, 2026 | Downloadable |
| Context | 1M | Large | Large |
| Terminal-Bench 2.1 | 82.7 | Competitive | Competitive |
| Best at | Cheapest hosted API | Self-host frontier-adjacent | Self-host general coding |
*DeepSeek doubles during peak hours.
Where Each Wins
- DeepSeek V4 Flash 0731 → cheapest managed API. Native Codex + Responses API support, DSpark speculative decoding, 82.7 Terminal-Bench 2.1. Almost always the cheapest way to run a capable coding agent.
- Kimi K3 → self-host frontier-adjacent. Weights opened July 27, 2026; flat $3/$15 if you use the hosted API. Strong choice when you want control and no peak/off-peak pricing games.
- GLM-5.2 → self-host general coding. Downloadable weights and solid all-round coding make it a dependable on-prem option.
Hosted API vs Self-Host
- Cheapest bill, no ops → DeepSeek V4 Flash 0731 hosted API
- Data control / on-prem → Kimi K3 or GLM-5.2 weights on your own GPUs
- Predictable flat pricing → Kimi K3
Self-hosting trades per-token fees for GPU cost and ops overhead — worth it at high volume or under strict data-residency rules.
Verdict
- Cheapest capable open-weight API → DeepSeek V4 Flash 0731
- Best self-host, flat pricing → Kimi K3
- Best self-host general coding → GLM-5.2
Sources
- DeepSeek API pricing: api-docs.deepseek.com/quick_start/pricing
- Moonshot AI — Kimi: moonshot.ai