Atria Dawn vs DeepSeek V4 Pro vs GLM-5.3 vs Kimi K3 (2026)
The short answer
Atria Dawn Preview for tool-heavy research agents, Kimi K3 or GLM-5.3 for coding, DeepSeek V4 Pro for cost and context. All four are open-weight Chinese models released or updated between July and September 2026 and all four post-train for agentic work rather than chat. Verified September 16, 2026.
| Atria Dawn Preview | DeepSeek V4 Pro (0813) | GLM-5.3 | Kimi K3 | |
|---|---|---|---|---|
| Developer | Shanghai AI Laboratory | DeepSeek | Z.ai (Zhipu) | Moonshot AI |
| Released | Sep 11, 2026 (weights); paper Sep 14 | GA Aug 13, 2026; repriced Aug 16 | Aug 14, 2026 | Jul 2026; open weights Jul 27 |
| Size | 744B MoE (GLM-5.2 base) | 1.6T MoE | 743B MoE (GLM-5.2 base) | Not fully disclosed; ~1T-class MoE |
| License | MIT | MIT | Open weights staged; not all shipped at launch | Modified MIT |
| Context | 256K | 1M (384K max output) | 1M | 1M |
| Input | Text only | Text + image (V4 Flash Vision line) | Text + image | Text + image |
| API price / MTok | Free preview | $0.66 / $1.98 off-peak; $1.32 / $3.96 peak | $1.40 / $4.40 | $3 / $15 flat |
| Cache hit | n/a | $0.022 off-peak | $0.26 | Published, cheaper than input |
Prices from vendor pages as recorded in our pricing reference, last verified September 4, 2026; DeepSeek quotes are off-peak, the correct default for US/EU working hours.
Benchmarks side by side (vendor-reported)
Shanghai AI Lab’s September 14 report is the only table that includes all four on the same benchmarks. Bold is the best of the four.
| Benchmark | Atria Dawn | DeepSeek V4 Pro | GLM-5.3 | Kimi K3 |
|---|---|---|---|---|
| BFCL v4 (tool calling) | 77.0 | 71.4 | 74.1 | 69.1 |
| AutomationBench | 53.8 | 41.7 | 49.2 | 45.9 |
| SkillsBench | 66.4 | 65.0 | 63.3 | 51.9 |
| τ³-Bench Banking | 41.2 | 44.3 | 40.2 | 37.1 |
| DeepSearchQA | 96.0 | – | 94.7 | 95.9 |
| BrowseComp | 92.5 | 83.4 | – | 91.2 |
| DeepResearch Bench II | 51.1 | 46.6 | 52.7 | 51.3 |
| SWE-bench Pro | 59.6 | 58.3 | 60.3 | 61.6 |
| Terminal-Bench 2.1 | 78.3 | 78.7 | 85.4 | – |
| MLE-bench Lite | 86.2 | 86.8 | 80.8 | 85.8 |
| Workspace-Bench | 65.0 | 55.7 | 63.9 | 60.6 |
| GDPval | 1583 | 1517 | 1667 | 1611 |
| JobBench | 50.3 | 54.1 | 58.2 | 54.3 |
| CyberGym | 86.5 | 83.3 | 84.5 | 78.7 |
Read it as a shape, not a ranking. Atria wins the agent-plumbing rows (function calling, automation, skills, search) and cybersecurity. GLM-5.3 wins the office-work rows (GDPval, JobBench) and terminal work. Kimi K3 edges SWE-bench Pro. DeepSeek V4 Pro rarely wins a row but is the cheapest by a wide margin.
Where each one fits
Atria Dawn Preview — research and scientific automation
Built for “research question → verifiable result” loops: plan, call tools, run code, analyze, recover. The paper includes a 769-task human study in which participants judged about a third of completed tasks infeasible without the model. Text-only input and a preview label are the limits. Details: What is Atria Dawn Preview?
DeepSeek V4 Pro — long-horizon agents on a budget
1M context, 384K max output, cache hits at $0.022 per MTok off-peak, concurrency 500. In our measured-cost reference a 30K-in/5K-out task costs about $0.03 off-peak against $0.064 for GLM-5.3 and $0.165 for Kimi K3. The trade is peak-hour doubling (01:00-04:00 and 06:00-10:00 UTC) and middling benchmark wins. See DeepSeek V4.1 Flash vs V4 Pro vs Gemini 3.8 Flash.
GLM-5.3 — office work and terminals
The strongest of the four on GDPval, JobBench and Terminal-Bench 2.1, with mandatory thinking and a $18/month Coding Plan for individuals. Atria is post-trained on GLM-5.3’s predecessor, GLM-5.2, so the two share a base and a footprint; GLM-5.3 is the more general model, Atria the more tool-specialized. See GLM-5.3 Flash vs Gemini 3.8 Flash vs GPT-5.6 Luna.
Kimi K3 — coding first
Best SWE-bench Pro of the four and the highest Artificial Analysis Intelligence Index among them (59.6, versus GLM-5.3 at 59.4; Atria and V4 Pro unscored or lower). Also the most expensive at $3/$15 flat, five times V4 Pro’s off-peak input price, which only pays off if the task fails on the cheaper models.
Decision table
| You want… | Pick |
|---|---|
| Deep research agent with many tool calls | Atria Dawn Preview (free preview), fall back to Kimi K3 |
| Cheapest acceptable agent at scale | DeepSeek V4 Pro off-peak |
| Coding agent in a terminal | GLM-5.3, then Kimi K3 |
| Office deliverables (decks, reports) | GLM-5.3 |
| Longest context / biggest outputs | DeepSeek V4 Pro (1M / 384K); GLM-5.3 and Kimi K3 also 1M |
| Cleanest license for redistribution | Atria or DeepSeek (MIT) |
| Image input required | Not Atria; any of the other three |
Caveats that apply to all four
- Vendor-reported scores. Only Kimi K3 and GLM-5.3 have Artificial Analysis Intelligence Index entries; Atria has none. Wait for audited scores before betting a product on a row.
- Token efficiency beats list price. Measured cost per task can differ from price-sheet arithmetic by 5-8x depending on how many tokens a model burns; see Token efficiency vs token price.
- Hosted endpoints are China-based for Atria (international endpoint aside), DeepSeek, Z.ai and Moonshot. For regulated data, self-host or use a Western reseller; see Stealth model vs named provider data risk.
- Frontier gap. Claude Opus 5 leads every coding row here, and Fable 5.1 leads the intelligence index at 65.6. These four are cost and openness choices.
Related
- What is Atria Dawn Preview?
- K2 Horizon 375B vs GLM-5.2 vs DeepSeek V4 Pro (September 2026)
- Best open-weight frontier model 2026, ranked
- DeepSeek vs Moonshot vs Z.ai vs MiniMax: the Chinese AI IPO race