V4-Pro vs GLM-5.3 vs Kimi K3: Cheap Frontier 2026
The Short Answer
DeepSeek V4-Pro still wins on price off-peak. GLM-5.3 wins on coding. Kimi K3 wins if you need the weights.
| Model | Input / Output (per MTok) | 30K-in / 5K-out | Open weights? |
|---|---|---|---|
| DeepSeek V4-Pro (off-peak) | $0.66 / $1.98 | $0.0297 | No |
| DeepSeek V4-Pro (peak) | $1.32 / $3.96 | $0.0594 | No |
| GLM-5.3 | $1.40 / $4.40 | $0.0640 | Staged, not at launch |
| Kimi K3 | $3 / $15 | $0.1650 | Yes (July 27, 2026) |
Verified August 17, 2026. DeepSeek’s peak/off-peak schedule took effect 16:00 UTC August 16, 2026.
The August 16 repricing cut DeepSeek’s lead from roughly 4x to 2x off-peak, and to near-parity at peak. For the first time in this category, price alone no longer settles the question.
What Changed
DeepSeek V4-Pro was a flat $0.435 / $0.87. It is now $0.66 / $1.98 off-peak and $1.32 / $3.96 at peak — an output increase of 127% to 355%. Cache-hit input rose from $0.003625 to $0.044 at peak, a 12x jump and the largest single change in the schedule.
Peak hours are 01:00-04:00 and 06:00-10:00 UTC. The remaining 17 hours are half price.
Where Each Wins
DeepSeek V4-Pro (GA August 13, 2026) is still the cheapest frontier-class option if you can schedule around the clock. It carries a 1M context and a 384K max output — the largest output ceiling of the three by a wide margin. Flexible reasoning effort (low / high / max) lets you buy thinking only where it pays, which is a genuine cost lever rather than a marketing bullet. Native OpenAI Responses API support with one-click Codex setup, plus an Anthropic-format endpoint, make it the easiest to drop into existing tooling. Concurrency is capped at 500.
GLM-5.3 (Z.ai, released August 14, 2026) is the coding specialist. It shares the 743B base with GLM-5.2 but is post-trained for agentic coding, and thinking mode is mandatory — you cannot turn it off, which is a deliberate quality-over-cost stance. Cached input is $0.26. The Coding Plan starts at $18/month, and off-peak usage (outside 14:00-18:00 UTC+8 on weekdays) consumes 50% fewer points. If your workload is a coding agent running all day, the subscription math beats per-token billing on all three.
Kimi K3 is the expensive one at $3/$15 flat — 5x V4-Pro off-peak — and it is the only one you can actually run yourself. Open weights landed July 27, 2026. Flat pricing with no peak/off-peak and no thinking-mode surcharge makes it the most predictable of the three. If your constraint is data residency, air-gapped deployment, or long-term independence from a vendor’s pricing decisions, the premium is the point.
The Clock Problem
DeepSeek’s advantage is now conditional on when you run. Peak in local time:
- US Eastern: 21:00-00:00, 02:00-06:00 — the workday is entirely off-peak.
- Central Europe: 03:00-06:00, 08:00-12:00 — mornings peak, afternoons free.
- China (UTC+8): 09:00-12:00, 14:00-18:00 — the whole business day.
GLM-5.3 has its own inverted schedule: its off-peak is outside 14:00-18:00 UTC+8 weekdays, which is a much narrower peak window than DeepSeek’s.
For a North American team, DeepSeek’s price increase is largely theoretical and it remains the clear cost winner. For an Asian team on business hours, V4-Pro peak ($0.0594) versus GLM-5.3 flat ($0.0640) is a 7% difference — at which point you should pick on capability, not price.
The Cache Consideration
If you run a large stable system prompt, cache-hit rates dominate:
| Model | Cache-hit input | vs base |
|---|---|---|
| DeepSeek V4-Pro (off-peak) | $0.022 | ~97% off |
| DeepSeek V4-Pro (peak) | $0.044 | ~97% off |
| V4-Pro (before Aug 16) | $0.003625 | ~99% off |
| GLM-5.3 | $0.26 | ~81% off |
DeepSeek still has the deepest cache discount by a wide margin — but its absolute cache cost rose 6-12x. Teams whose architecture was built on near-free cache hits took the largest increase in this comparison, even though the headline percentages looked worse elsewhere.
The Decision Framework
- Cheapest frontier tokens, Western working hours → DeepSeek V4-Pro. Off-peak pricing is still roughly half of GLM-5.3.
- Sustained coding agent, want a flat subscription → GLM-5.3 Coding Plan at $18/month. Per-token billing loses to subscriptions on continuous agent work.
- Need the weights — residency, air-gap, or independence → Kimi K3. It is the only real option, and the 5x premium buys optionality.
- Asia-Pacific business hours → GLM-5.3. You’d pay DeepSeek’s peak rate most of the day for a 7% saving.
- Very long outputs → DeepSeek V4-Pro, at 384K max output.
- Want one bill you can forecast without a scheduler → Kimi K3. Flat, no clock, no thinking surcharge.
Last verified: August 17, 2026. Prices from official vendor pricing pages.