AI agents · OpenClaw · self-hosting · automation

Quick Answer

DeepSeek V4 vs Kimi K3 vs Inkling: Open-Weight (Jul 2026)

Published:

DeepSeek V4 vs Kimi K3 vs Inkling: Open-Weight (Jul 2026)

Three frontier-adjacent open-weight models, three different bets on how open AI should look. As of July 20, 2026:

  • DeepSeek V4-Pro — MIT-licensed, 1.6T params, GA release imminent this week
  • Kimi K3 — 2.8T-param frontier contender, weights dropping July 27
  • Inkling — Apache 2.0, Thinking Machines (Mira Murati’s startup), frontier-adjacent

Together they represent the strongest open-weight frontier month in AI history. This comparison covers license, benchmarks, self-hosting economics, and ecosystem considerations for enterprise decision-making.

Last verified: July 20, 2026

Head-to-Head Table

FeatureDeepSeek V4-ProKimi K3Inkling
DeveloperDeepSeekMoonshot AIThinking Machines
FounderLiang WenfengYang ZhilinMira Murati (ex-OpenAI CTO)
CountryChinaChinaUnited States
Preview / release dateApril 24, 2026 (V4 Preview); GA imminentJuly 16, 2026 (API); July 27, 2026 (weights)July 2026
Parameters (total)1.6 trillion2.8 trillionUndisclosed
Parameters (active)49 billion~50-100B effective (16 of 896 experts active)Undisclosed
LicenseMITOpen weights (specific terms at July 27)Apache 2.0
Context window1M tokens1M tokensUndisclosed but frontier-competitive
Attention mechanismCompressed Sparse Attention (CSA)Kimi Delta Attention (KDA)Undisclosed
API pricing~$0.30 / $1.20 per MTok (expected, V4 pricing)$3 / $15 per MTokNot yet public API
Aggregate leaderboard~#7-8 (V4 Preview)#4 (80.96)Frontier-adjacent
Strongest atReasoning benchmarks, long context, extreme cost efficiencyFrontier-adjacent overall, #1 Frontend Code arenaAcademic benchmarks, agent workflows
Self-host GPU need (est.)4-8 H100/H2008-16 H100/H2004-12 H100/H200 (est.)
Weights on HuggingFaceYes (V4 Preview)Coming July 27Yes

What Each Model Is

DeepSeek V4-Pro — The MIT-Licensed Efficiency Leader

Positioning: DeepSeek’s flagship open-weight model. V4 Preview shipped April 24, 2026 under MIT license — the most permissive license in mainstream AI. GA release is imminent (within days of July 20) after DeepSeek committed to a mid-July target and multiple industry publications confirm the launch window is closing this week.

Key specs:

  • 1.6T total parameters, 49B active — mixture-of-experts architecture keeps inference cost tractable.
  • Compressed Sparse Attention (CSA) — DeepSeek’s novel long-context attention mechanism.
  • 1M-token default context — matches Kimi K3, exceeds Sol’s 400K.
  • MIT license — unlimited commercial use, no copyleft, no restrictions.

Ecosystem signal: DeepSeek raised $7.4B in its first external funding round in July 2026, aimed at global expansion. This is the strongest signal yet that DeepSeek is committing to the open-weight strategy long-term.

Best for: Enterprise self-hosting for compliance / cost, large-scale inference workloads, applications where MIT license simplifies legal review.

Kimi K3 — The Benchmark-Leading Open-Weight Model

Positioning: Moonshot AI’s frontier-adjacent bet. Released via API on July 16, 2026; open weights drop July 27, 2026. #4 on public benchmark leaderboards behind only proprietary Fable 5, Sol, and Gemini 3.5 Pro.

Key specs:

  • 2.8 trillion total parameters — largest of the three.
  • Stable LatentMoE with 896 experts, 16 active per token — high specialization, tractable inference.
  • Kimi Delta Attention (KDA) — hybrid linear + Attention Residuals for long-context efficiency.
  • Native multimodal — image + text handling.
  • 1M-token context.
  • MXFP4 quantization ready — dramatically reduced memory footprint for deployment.

Ecosystem signal: #1 on Arena.ai’s Frontend Code arena — the best coding benchmark in one specific arena, ahead of both proprietary Sol and Fable 5. Strong demonstration that open-weight can lead specific verticals.

Best for: Coding workflows (especially frontend), general-purpose applications where benchmark leadership matters, self-hosting for cost + compliance in Q4 2026.

Inkling — The Thinking Machines Open-Weight Play

Positioning: Thinking Machines is Mira Murati’s post-OpenAI startup — she’s the former OpenAI CTO. Inkling is the company’s first open-weight release, positioned as Apache 2.0 to signal maximum enterprise-friendliness.

Key specs:

  • Apache 2.0 license — near-MIT permissiveness with patent-grant provisions (arguably better for enterprise commercial deployment).
  • Frontier-adjacent benchmarks — competitive with V4-Pro Preview on academic tests.
  • Parameter count undisclosed publicly (as of July 20, 2026).
  • US-origin — different geopolitical positioning than Chinese-origin V4 and K3.

Ecosystem signal: Mira Murati’s involvement is the story. Her post-OpenAI positioning as a US-origin, open-weight competitor to closed OpenAI is deliberate. Inkling is Thinking Machines’ opening statement.

Best for: Enterprise deployments where US-origin matters, applications that value Apache 2.0’s patent-grant provisions, users prioritizing Mira Murati’s track record over pure benchmark leadership.

Head-to-Head on Key Dimensions

License Permissiveness

ModelLicenseCommercial useModificationRedistributionNotes
DeepSeek V4-ProMITMost permissive; no restrictions
InklingApache 2.0Near-MIT with patent grants; strong for commercial
Kimi K3Open weights (specific terms at July 27)Likely ✓Likely ✓Likely with attributionModified license; review terms carefully

Winner: DeepSeek V4-Pro (MIT). Runner-up: Inkling (Apache 2.0). Kimi K3 competitive pending specific license terms.

Benchmark Performance

ModelAggregate scoreBest specific benchmark
Kimi K380.96 (#4 overall)#1 Arena.ai Frontend Code
DeepSeek V4-Pro Preview~78Strong on reasoning, long-context
InklingFrontier-adjacent (varies)Academic benchmarks

Winner: Kimi K3 on aggregate. DeepSeek V4-Pro GA release may shift this — expected refined weights close gaps.

API + Access Availability

ModelAPI available todayWeights available todayEnterprise via cloud
DeepSeek V4-Pro✓ (V4 Preview)✓ (V4 Preview)Via DeepSeek + third-party
Kimi K3Coming July 27Cloudflare Workers AI, OpenRouter
InklingLimitedThird-party hosts emerging

Winner: DeepSeek V4-Pro — most mature availability. Kimi K3 close second — API available, weights imminent. Inkling less broadly available.

Self-Hosting Economics

Rough monthly self-hosted cost estimates (medium-throughput deployment on cloud H100/H200s):

ModelEstimated monthly costEffective active paramsNotes
DeepSeek V4-Pro$1500-300049BBest efficiency; MoE keeps inference cheap
Inkling$2000-4000 (est.)UndisclosedDetails pending broader deployment
Kimi K3$2500-5000Higher effective compute2.8T total demands more infrastructure

Winner: DeepSeek V4-Pro — lowest self-hosting compute cost. Kimi K3 more expensive but wins on capability-per-dollar for high-capability workloads.

Ecosystem + Community

ModelHuggingFace ecosystemFine-tuning supportCommunity distills
DeepSeek V4-ProStrong (multiple V3/V4 forks)YesExtensive (100+ derivatives on V3)
Kimi K3Growing (July 27 weight release)Expected yesNascent
InklingEmerging (Thinking Machines’ first release)YesNascent

Winner: DeepSeek V4-Pro by ecosystem maturity — DeepSeek V3 had massive community adoption; V4-Pro inherits that.

Real Use Case Comparisons

Use Case 1: Enterprise Self-Hosting for Compliance

Winner: DeepSeek V4-Pro (MIT license). Zero-restriction commercial deployment; strongest self-hosting economics. Runner-up: Inkling (Apache 2.0). Runner-up US-origin option.

Use Case 2: Cost-Optimized Large-Scale Inference

Winner: DeepSeek V4-Pro — API pricing extrapolation ~$0.30/$1.20 per MTok undercuts Kimi K3 by ~10x and Sol by ~15x. Runner-up: Kimi K3 at $3/$15 per MTok.

Use Case 3: Frontend Code Generation

Winner: Kimi K3 — #1 on Arena.ai Frontend Code arena. No competition here among open-weight models.

Use Case 4: Fine-Tuning for Domain-Specific Model

Winner: DeepSeek V4-Pro — most mature fine-tuning tooling and community examples. Runner-up: Kimi K3 post-July 27 as tooling matures.

Use Case 5: US-Origin Preference (Compliance / Geopolitics)

Winner: Inkling — the only US-origin option of the three. Note: open weights + on-prem deployment mitigates most geopolitical concerns with V4-Pro and K3 (data stays in your infrastructure).

Use Case 6: Long-Context Deep Analysis (500K+ tokens)

Tie: V4-Pro or K3 — both 1M context. V4-Pro slightly more efficient at long context due to CSA architecture; K3 has more raw capability. Test both on your workload.

Use Case 7: Academic Research / Reproducibility

Winner: Inkling (Apache 2.0) for maximum reuse rights. Runner-up: DeepSeek V4-Pro (MIT).

Use Case 8: Latest-Capability at Any Cost

Winner: Kimi K3 — highest aggregate benchmark score of the three. Watch: V4 GA release could shift this within days.

The Real Decision Framework

Pick DeepSeek V4-Pro if:

  • MIT license simplifies legal review.
  • Self-hosting economics dominate.
  • You want to inherit DeepSeek V3’s mature ecosystem.
  • Cheapest API pricing extrapolation.

Pick Kimi K3 if:

  • Frontend code generation is core.
  • You want highest aggregate benchmark open-weight model.
  • 2.8T param scale matters for your workload.
  • Comfortable with Chinese-origin model post-open-weights.

Pick Inkling if:

  • US-origin matters for compliance.
  • Apache 2.0’s patent-grant provisions matter.
  • Mira Murati’s post-OpenAI track record is meaningful to you.
  • You’re okay with less mature ecosystem than V4-Pro.

Use two or three in parallel:

  • DeepSeek V4-Pro for high-volume inference + K3 for hardest coding tasks.
  • Inkling for US-origin regulatory workloads + V4-Pro for cost-optimized batch work.
  • K3 for benchmark-critical tasks + V4-Pro for cost-critical tasks.

Enterprises with $50K+ AI budgets frequently run two of the three for workload-appropriate routing.

The Bigger Picture: Q3 2026 Open-Weight Consolidation

Three frontier-adjacent open-weight releases in a two-week window is unprecedented. Combined with:

  • MiniMax M3 Pro (2.7T open weights, released mid-July)
  • GLM-5.2 (Zhipu, open weights, below frontier)
  • openPangu 2.0 (Huawei, open weights, below frontier)

The July 2026 open-weight wave establishes that frontier-adjacent AI capability is no longer proprietary-only. Enterprises have real strategic choice on how to build their AI stacks — and open-weight is now a first-class option, not a fallback.

Expect:

  1. Consolidation to 3-5 dominant open-weight models by end of 2026 (V4-Pro, K3, Inkling likely among them).
  2. Community distillations proliferating — 100+ specialized variants on HuggingFace.
  3. Enterprise self-hosting infrastructure buildouts in Q4 2026 as compliance-motivated deployments accelerate.
  4. Proprietary model pricing pressure — Sol and Fable 5 will need to justify their premium over $3/$15 K3 API and effectively-free self-hosted V4-Pro.

Bottom Line

No universal winner in July 2026 — each open-weight model wins specific categories.

  • DeepSeek V4-Pro — best MIT license, most mature ecosystem, cheapest self-hosting economics.
  • Kimi K3 — best benchmarks, best frontend code, largest parameter count.
  • Inkling — best Apache 2.0 license, best US-origin option, backed by Mira Murati’s track record.

For most enterprises evaluating this week: wait for DeepSeek V4 GA (imminent, days away) and Kimi K3 weights (July 27). Benchmark both against your workload, then pick primary + fallback. Don’t lock in a single-model strategy — the open-weight frontier is too fluid right now.

For developers exploring: try Kimi K3 API today ($3/$15 per MTok is cheap enough for real experimentation). Move to self-hosting evaluation after July 27 weights release.

For strategy leads: the open-weight frontier is real capability in 2026. Building AI stack strategy without considering V4-Pro, K3, and Inkling as first-class options is a mistake.

Sources