DeepSeek V4 vs Kimi K3 vs Inkling: Open-Weight (Jul 2026)
DeepSeek V4 vs Kimi K3 vs Inkling: Open-Weight (Jul 2026)
Three frontier-adjacent open-weight models, three different bets on how open AI should look. As of July 20, 2026:
- DeepSeek V4-Pro — MIT-licensed, 1.6T params, GA release imminent this week
- Kimi K3 — 2.8T-param frontier contender, weights dropping July 27
- Inkling — Apache 2.0, Thinking Machines (Mira Murati’s startup), frontier-adjacent
Together they represent the strongest open-weight frontier month in AI history. This comparison covers license, benchmarks, self-hosting economics, and ecosystem considerations for enterprise decision-making.
Last verified: July 20, 2026
Head-to-Head Table
| Feature | DeepSeek V4-Pro | Kimi K3 | Inkling |
|---|---|---|---|
| Developer | DeepSeek | Moonshot AI | Thinking Machines |
| Founder | Liang Wenfeng | Yang Zhilin | Mira Murati (ex-OpenAI CTO) |
| Country | China | China | United States |
| Preview / release date | April 24, 2026 (V4 Preview); GA imminent | July 16, 2026 (API); July 27, 2026 (weights) | July 2026 |
| Parameters (total) | 1.6 trillion | 2.8 trillion | Undisclosed |
| Parameters (active) | 49 billion | ~50-100B effective (16 of 896 experts active) | Undisclosed |
| License | MIT | Open weights (specific terms at July 27) | Apache 2.0 |
| Context window | 1M tokens | 1M tokens | Undisclosed but frontier-competitive |
| Attention mechanism | Compressed Sparse Attention (CSA) | Kimi Delta Attention (KDA) | Undisclosed |
| API pricing | ~$0.30 / $1.20 per MTok (expected, V4 pricing) | $3 / $15 per MTok | Not yet public API |
| Aggregate leaderboard | ~#7-8 (V4 Preview) | #4 (80.96) | Frontier-adjacent |
| Strongest at | Reasoning benchmarks, long context, extreme cost efficiency | Frontier-adjacent overall, #1 Frontend Code arena | Academic benchmarks, agent workflows |
| Self-host GPU need (est.) | 4-8 H100/H200 | 8-16 H100/H200 | 4-12 H100/H200 (est.) |
| Weights on HuggingFace | Yes (V4 Preview) | Coming July 27 | Yes |
What Each Model Is
DeepSeek V4-Pro — The MIT-Licensed Efficiency Leader
Positioning: DeepSeek’s flagship open-weight model. V4 Preview shipped April 24, 2026 under MIT license — the most permissive license in mainstream AI. GA release is imminent (within days of July 20) after DeepSeek committed to a mid-July target and multiple industry publications confirm the launch window is closing this week.
Key specs:
- 1.6T total parameters, 49B active — mixture-of-experts architecture keeps inference cost tractable.
- Compressed Sparse Attention (CSA) — DeepSeek’s novel long-context attention mechanism.
- 1M-token default context — matches Kimi K3, exceeds Sol’s 400K.
- MIT license — unlimited commercial use, no copyleft, no restrictions.
Ecosystem signal: DeepSeek raised $7.4B in its first external funding round in July 2026, aimed at global expansion. This is the strongest signal yet that DeepSeek is committing to the open-weight strategy long-term.
Best for: Enterprise self-hosting for compliance / cost, large-scale inference workloads, applications where MIT license simplifies legal review.
Kimi K3 — The Benchmark-Leading Open-Weight Model
Positioning: Moonshot AI’s frontier-adjacent bet. Released via API on July 16, 2026; open weights drop July 27, 2026. #4 on public benchmark leaderboards behind only proprietary Fable 5, Sol, and Gemini 3.5 Pro.
Key specs:
- 2.8 trillion total parameters — largest of the three.
- Stable LatentMoE with 896 experts, 16 active per token — high specialization, tractable inference.
- Kimi Delta Attention (KDA) — hybrid linear + Attention Residuals for long-context efficiency.
- Native multimodal — image + text handling.
- 1M-token context.
- MXFP4 quantization ready — dramatically reduced memory footprint for deployment.
Ecosystem signal: #1 on Arena.ai’s Frontend Code arena — the best coding benchmark in one specific arena, ahead of both proprietary Sol and Fable 5. Strong demonstration that open-weight can lead specific verticals.
Best for: Coding workflows (especially frontend), general-purpose applications where benchmark leadership matters, self-hosting for cost + compliance in Q4 2026.
Inkling — The Thinking Machines Open-Weight Play
Positioning: Thinking Machines is Mira Murati’s post-OpenAI startup — she’s the former OpenAI CTO. Inkling is the company’s first open-weight release, positioned as Apache 2.0 to signal maximum enterprise-friendliness.
Key specs:
- Apache 2.0 license — near-MIT permissiveness with patent-grant provisions (arguably better for enterprise commercial deployment).
- Frontier-adjacent benchmarks — competitive with V4-Pro Preview on academic tests.
- Parameter count undisclosed publicly (as of July 20, 2026).
- US-origin — different geopolitical positioning than Chinese-origin V4 and K3.
Ecosystem signal: Mira Murati’s involvement is the story. Her post-OpenAI positioning as a US-origin, open-weight competitor to closed OpenAI is deliberate. Inkling is Thinking Machines’ opening statement.
Best for: Enterprise deployments where US-origin matters, applications that value Apache 2.0’s patent-grant provisions, users prioritizing Mira Murati’s track record over pure benchmark leadership.
Head-to-Head on Key Dimensions
License Permissiveness
| Model | License | Commercial use | Modification | Redistribution | Notes |
|---|---|---|---|---|---|
| DeepSeek V4-Pro | MIT | ✓ | ✓ | ✓ | Most permissive; no restrictions |
| Inkling | Apache 2.0 | ✓ | ✓ | ✓ | Near-MIT with patent grants; strong for commercial |
| Kimi K3 | Open weights (specific terms at July 27) | Likely ✓ | Likely ✓ | Likely with attribution | Modified license; review terms carefully |
Winner: DeepSeek V4-Pro (MIT). Runner-up: Inkling (Apache 2.0). Kimi K3 competitive pending specific license terms.
Benchmark Performance
| Model | Aggregate score | Best specific benchmark |
|---|---|---|
| Kimi K3 | 80.96 (#4 overall) | #1 Arena.ai Frontend Code |
| DeepSeek V4-Pro Preview | ~78 | Strong on reasoning, long-context |
| Inkling | Frontier-adjacent (varies) | Academic benchmarks |
Winner: Kimi K3 on aggregate. DeepSeek V4-Pro GA release may shift this — expected refined weights close gaps.
API + Access Availability
| Model | API available today | Weights available today | Enterprise via cloud |
|---|---|---|---|
| DeepSeek V4-Pro | ✓ (V4 Preview) | ✓ (V4 Preview) | Via DeepSeek + third-party |
| Kimi K3 | ✓ | Coming July 27 | Cloudflare Workers AI, OpenRouter |
| Inkling | Limited | ✓ | Third-party hosts emerging |
Winner: DeepSeek V4-Pro — most mature availability. Kimi K3 close second — API available, weights imminent. Inkling less broadly available.
Self-Hosting Economics
Rough monthly self-hosted cost estimates (medium-throughput deployment on cloud H100/H200s):
| Model | Estimated monthly cost | Effective active params | Notes |
|---|---|---|---|
| DeepSeek V4-Pro | $1500-3000 | 49B | Best efficiency; MoE keeps inference cheap |
| Inkling | $2000-4000 (est.) | Undisclosed | Details pending broader deployment |
| Kimi K3 | $2500-5000 | Higher effective compute | 2.8T total demands more infrastructure |
Winner: DeepSeek V4-Pro — lowest self-hosting compute cost. Kimi K3 more expensive but wins on capability-per-dollar for high-capability workloads.
Ecosystem + Community
| Model | HuggingFace ecosystem | Fine-tuning support | Community distills |
|---|---|---|---|
| DeepSeek V4-Pro | Strong (multiple V3/V4 forks) | Yes | Extensive (100+ derivatives on V3) |
| Kimi K3 | Growing (July 27 weight release) | Expected yes | Nascent |
| Inkling | Emerging (Thinking Machines’ first release) | Yes | Nascent |
Winner: DeepSeek V4-Pro by ecosystem maturity — DeepSeek V3 had massive community adoption; V4-Pro inherits that.
Real Use Case Comparisons
Use Case 1: Enterprise Self-Hosting for Compliance
Winner: DeepSeek V4-Pro (MIT license). Zero-restriction commercial deployment; strongest self-hosting economics. Runner-up: Inkling (Apache 2.0). Runner-up US-origin option.
Use Case 2: Cost-Optimized Large-Scale Inference
Winner: DeepSeek V4-Pro — API pricing extrapolation ~$0.30/$1.20 per MTok undercuts Kimi K3 by ~10x and Sol by ~15x. Runner-up: Kimi K3 at $3/$15 per MTok.
Use Case 3: Frontend Code Generation
Winner: Kimi K3 — #1 on Arena.ai Frontend Code arena. No competition here among open-weight models.
Use Case 4: Fine-Tuning for Domain-Specific Model
Winner: DeepSeek V4-Pro — most mature fine-tuning tooling and community examples. Runner-up: Kimi K3 post-July 27 as tooling matures.
Use Case 5: US-Origin Preference (Compliance / Geopolitics)
Winner: Inkling — the only US-origin option of the three. Note: open weights + on-prem deployment mitigates most geopolitical concerns with V4-Pro and K3 (data stays in your infrastructure).
Use Case 6: Long-Context Deep Analysis (500K+ tokens)
Tie: V4-Pro or K3 — both 1M context. V4-Pro slightly more efficient at long context due to CSA architecture; K3 has more raw capability. Test both on your workload.
Use Case 7: Academic Research / Reproducibility
Winner: Inkling (Apache 2.0) for maximum reuse rights. Runner-up: DeepSeek V4-Pro (MIT).
Use Case 8: Latest-Capability at Any Cost
Winner: Kimi K3 — highest aggregate benchmark score of the three. Watch: V4 GA release could shift this within days.
The Real Decision Framework
Pick DeepSeek V4-Pro if:
- MIT license simplifies legal review.
- Self-hosting economics dominate.
- You want to inherit DeepSeek V3’s mature ecosystem.
- Cheapest API pricing extrapolation.
Pick Kimi K3 if:
- Frontend code generation is core.
- You want highest aggregate benchmark open-weight model.
- 2.8T param scale matters for your workload.
- Comfortable with Chinese-origin model post-open-weights.
Pick Inkling if:
- US-origin matters for compliance.
- Apache 2.0’s patent-grant provisions matter.
- Mira Murati’s post-OpenAI track record is meaningful to you.
- You’re okay with less mature ecosystem than V4-Pro.
Use two or three in parallel:
- DeepSeek V4-Pro for high-volume inference + K3 for hardest coding tasks.
- Inkling for US-origin regulatory workloads + V4-Pro for cost-optimized batch work.
- K3 for benchmark-critical tasks + V4-Pro for cost-critical tasks.
Enterprises with $50K+ AI budgets frequently run two of the three for workload-appropriate routing.
The Bigger Picture: Q3 2026 Open-Weight Consolidation
Three frontier-adjacent open-weight releases in a two-week window is unprecedented. Combined with:
- MiniMax M3 Pro (2.7T open weights, released mid-July)
- GLM-5.2 (Zhipu, open weights, below frontier)
- openPangu 2.0 (Huawei, open weights, below frontier)
The July 2026 open-weight wave establishes that frontier-adjacent AI capability is no longer proprietary-only. Enterprises have real strategic choice on how to build their AI stacks — and open-weight is now a first-class option, not a fallback.
Expect:
- Consolidation to 3-5 dominant open-weight models by end of 2026 (V4-Pro, K3, Inkling likely among them).
- Community distillations proliferating — 100+ specialized variants on HuggingFace.
- Enterprise self-hosting infrastructure buildouts in Q4 2026 as compliance-motivated deployments accelerate.
- Proprietary model pricing pressure — Sol and Fable 5 will need to justify their premium over $3/$15 K3 API and effectively-free self-hosted V4-Pro.
Bottom Line
No universal winner in July 2026 — each open-weight model wins specific categories.
- DeepSeek V4-Pro — best MIT license, most mature ecosystem, cheapest self-hosting economics.
- Kimi K3 — best benchmarks, best frontend code, largest parameter count.
- Inkling — best Apache 2.0 license, best US-origin option, backed by Mira Murati’s track record.
For most enterprises evaluating this week: wait for DeepSeek V4 GA (imminent, days away) and Kimi K3 weights (July 27). Benchmark both against your workload, then pick primary + fallback. Don’t lock in a single-model strategy — the open-weight frontier is too fluid right now.
For developers exploring: try Kimi K3 API today ($3/$15 per MTok is cheap enough for real experimentation). Move to self-hosting evaluation after July 27 weights release.
For strategy leads: the open-weight frontier is real capability in 2026. Building AI stack strategy without considering V4-Pro, K3, and Inkling as first-class options is a mistake.
Sources
- DeepSeek Wikipedia (V4 details): en.wikipedia.org/wiki/DeepSeek
- Kimi K3 official page: moonshot.ai
- Digital Applied July 2026 open-weight wave analysis: digitalapplied.com/blog/open-weight-model-wave-july-2026-momentum-tracker
- Simon Willison Kimi K3 analysis: simonwillison.net/2026/Jul/16/kimi-k3