AI agents · OpenClaw · self-hosting · automation

Quick Answer

What Is Inkling? Thinking Machines' Open Model (Jul 2026)

Published:

What Is Inkling? Thinking Machines’ Open Model (Jul 2026)

On July 15, 2026, Thinking Machines Lab — the AI startup founded by former OpenAI CTO Mira Murati — released Inkling, its first in-house model. It’s a 975B-parameter, 41B-active-parameter Mixture-of-Experts transformer with a 1M-token context window, released under Apache 2.0. Full weights are available on Hugging Face.

Inkling is the first serious US-based open-weight frontier model at a time when the leading open-weight models (Kimi K3, DeepSeek V4 Pro, GLM-5.2, Qwen) are almost entirely Chinese.

Last verified: July 19, 2026

The Basics

SpecInkling
VendorThinking Machines Lab
Release dateJuly 15, 2026
LicenseApache 2.0
ArchitectureMixture-of-Experts (MoE) decoder-only transformer
Parameters975B total / 41B active per token
Context window1,000,000 tokens
ModalityMultimodal (text + image; audio in the model card)
Reasoning modesControllable “reasoning effort” (low / medium / high)
Training tokens45T multimodal tokens
BF16 full weights~1.9 TB on disk
Dynamic GGUF (Unsloth)270–285 GB
Fine-tuning platformTinker (Thinking Machines’ own)
Runtime supportTransformers, SGLang, vLLM, TokenSpeed, Unsloth

What Is Thinking Machines Lab?

Founded in 2024 by former OpenAI CTO Mira Murati and a group of ex-OpenAI researchers. Raised a record $2B seed round at a $12B valuation in 2025 before releasing a product. Nvidia is among the investors and has since deepened the relationship. The company’s stated mission: “build AI that extends human will and judgment,” with the strategic bet being customization over one-size-fits-all frontier models.

Prior to Inkling, Thinking Machines had released:

  • Tinker — a customization / fine-tuning platform for existing open-weight models.
  • Interactive Collaboration preview — an AI system for real-time human-AI work (not a shipped product yet).
  • Novel research (particularly on determinism in LLM inference).

Inkling completes the story: Thinking Machines now has its own model at the base of the stack, not just a wrapper over other people’s weights.

What Makes Inkling Different

1. Controllable reasoning effort. Inkling exposes a “thinking effort” dial (low / medium / high) similar to how GPT-5.6 Sol/Terra/Luna offer variants, but as a runtime toggle on a single model. Low is faster and cheaper; high burns more tokens for harder problems.

2. Full Apache 2.0. No community-license restrictions, no acceptable-use carveouts beyond what Apache 2.0 permits. This is more permissive than Meta’s Llama Community License (which restricts services above 700M MAU) and matches the permissive licensing of Qwen and DeepSeek.

3. US-based supply chain. For enterprises with China-supply-chain restrictions (US federal, defense, some EU), Inkling is the first truly frontier-class open-weight option that isn’t Chinese. That’s a real procurement differentiator.

4. Native 1M context. Not a retrofit — designed for long-context from training, using speculative MTP layers for faster inference at the 1M scale.

5. Trained partly on Kimi K2.5 data. The Wikipedia entry (as of July 17, 2026) notes Inkling was “partially trained using data from Moonshot AI’s Kimi K2.5 model,” which is unusual disclosure. It signals Thinking Machines’ willingness to build on Chinese open-weight foundations, not compete on scratch training alone.

The Fine Print (Where Inkling Loses)

Benchmark intelligence. Inkling does not top the Artificial Analysis Intelligence Index. Kimi K3 (released two days later, on July 17, 2026) outscores Inkling on general intelligence. GPT-5.6 Sol and Claude Fable 5 remain ahead of both on frontier reasoning.

Coding leaderboards. Kimi K3 hit #1 on the community-voted Arena Frontend Code leaderboard, surpassing Claude Fable 5. Inkling did not have a comparable coding-benchmark headline at launch.

Ecosystem size. Meta’s Muse Spark and Llama-descended models still have the largest deployment ecosystem (adapters, quantizations, community fine-tunes). Inkling’s ecosystem is one week old.

Multimodal input types. Inkling accepts text and image inputs (with audio in the model card as a capability). Video and richer audio are not primary launch strengths — Qwen 3.5 and Gemini 3.1 Pro are stronger on those.

Sub-Questions People Are Asking

Who should use Inkling?

  • Enterprises needing a US-supply-chain open-weight model for regulated deployments where Chinese-origin models are procurement-blocked.
  • Researchers wanting a permissive base model for fine-tune studies.
  • Product teams using Thinking Machines’ Tinker platform (Inkling is the natural default there).
  • Anyone customizing for domain-specific use cases where GPT/Claude are overkill and Kimi K3’s Chinese origin is a blocker.

Who should skip it?

  • Teams needing the highest-scoring frontier model — use GPT-5.6 Sol or Claude Fable 5.
  • Teams prioritizing coding-agent throughput — Claude Sonnet 5 (default in Claude Code) is stronger.
  • Teams needing the largest ecosystem today — Muse Spark or Kimi K3 have more third-party support.

Is Inkling free? Weights are free (Apache 2.0). Inference costs whatever your infrastructure runs it on — from Thinking Machines’ hosted API, to self-hosted GPU clusters, to consumer-hardware GGUFs via Unsloth. Fine-tuning through Tinker has commercial pricing (not publicly listed at launch).

How does controllable reasoning effort work? A runtime parameter, similar to OpenAI’s reasoning-effort toggles. Low mode approximates the speed of a non-reasoning model; high mode extends chain-of-thought for harder tasks. This is one of the practical selling points versus running separate small/large model variants.

Will Thinking Machines release more open-weight models? Not committed publicly. But Inkling’s positioning as “our open-weights model” (versus “the open-weights model”) suggests a family — probably matched to the reasoning-effort dial (smaller Inkling variants at 200-400B total, and possibly a larger 2T+ variant).

How does the funding runway look? Reported $2B seed at $12B valuation in 2025 with rumored follow-on rounds in 2026 (unconfirmed as to size and pricing). Mira Murati’s declining OpenAI residuals and rising Thinking Machines equity make the funding structure unusual: high leverage on this single company’s success.

The Strategic Read

Inkling is Thinking Machines’ entry into a specific market: US-based open-weight for enterprises. The competitive frame:

CompetitorPositioning vs Inkling
Kimi K3 (Moonshot)Bigger, cheaper per token, Chinese origin (procurement risk in US federal/defense)
DeepSeek V4 ProCheapest at frontier open-weight; Chinese origin
GLM-5.2 (Zhipu)Strong Chinese enterprise; less international presence
Muse Spark (Meta)Larger ecosystem; community-license restrictions above 700M MAU
Qwen 3.5 (Alibaba)Excellent multilingual (Chinese-first); Apache 2.0
Grok 4.5 (xAI)Not open-weight; runs on X infrastructure
Claude Sonnet 5 / GPT-5.6Closed-weight; higher benchmarks; higher costs at scale

Inkling wins procurement-blocked deals for US-based Apache-2.0 open weights. It doesn’t win benchmark-topping deals. That’s a deliberate strategic choice — Murati is not trying to out-compute the frontier labs. She’s trying to be the default open-weight vendor for enterprises that can’t or won’t use Chinese models.

What This Means for the Market

Open-weight competition just got a second US challenger (after Meta / Muse Spark). That materially changes the strategic picture for procurement-restricted buyers.

Customization-as-a-service is a real market. Tinker + Inkling is the first serious integrated stack from a US vendor. Enterprises that fine-tune matter — not just infer.

Chinese open-weight remains benchmark-strongest. Kimi K3 leading the Arena Frontend Code board and matching Claude Fable 5 on hard benchmarks with permissive licensing is the actual news story of July 2026. Inkling’s role is to give US-restricted buyers an alternative — not to beat Kimi K3 on capability.

Nvidia gets deeper hooks. Inkling was trained on Nvidia infrastructure with Nvidia investment. Every fine-tuning workload on Tinker plausibly runs on Nvidia. That’s structural revenue.

Bottom Line

Inkling is the first US-based Apache 2.0 open-weight frontier-class model. It’s not the highest-scoring open-weight model (Kimi K3 wins that) or the most-deployed (Muse Spark wins that). It’s the model you deploy when you need frontier-class open weights AND you can’t or won’t use Chinese-origin models.

Mira Murati’s bet is that “customization over one-size-fits-all” is the durable enterprise thesis — and Inkling + Tinker is the vehicle. Six months of enterprise deployment data will tell us if that thesis is right.

For now: download the weights, run inference, and evaluate against your specific workload before committing to any open-weight vendor. The open-weight race has more contenders than at any prior point in the LLM era.

Sources