What Is Naive-N0.5-Flash? NaiveAI's 309B Open MoE (2026)
The short answer
Naive-N0.5-Flash is a 309B-parameter, 15.5B-active MoE model with a native 1M-token context and no full-attention layers, released with MIT-licensed weights on September 27, 2026 by NaiveAI. It targets coding and AI research-and-development workloads, was continued-pretrained from Xiaomi’s open MiMo-V2.5 base on 3.25T tokens, and is priced at $0.10 input / $0.40 output / $0.01 cache read per million tokens on NaiveAI’s API. The company’s claim to fame is that the research pipeline itself was “substantially executed by AI systems,” with humans setting objectives and evaluation standards.
Who is NaiveAI
NaiveAI is a Beijing startup led by Tsinghua University’s Jifeng Dai, reportedly valued at about $1.4 billion. Its tagline is “Building Frontier AI with AI”: the stated method is to have AI systems run most of the research and engineering loop — architecture search, kernel optimization, training decisions — under human-defined objectives. Naive-N0.5-Flash is its first open-weight release and the NaiveRT inference stack is presented as a case study of the same process.
Architecture in one table
| Property | Naive-N0.5-Flash |
|---|---|
| Type | Mixture-of-Experts |
| Total / active parameters | 309B / 15.5B |
| Base model | Xiaomi MiMo-V2.5 (open weights) |
| Context | Native 1M tokens |
| Layers | 48 (eight six-layer modules) |
| Attention layout | 39 sliding-window (128-token window) + 9 DeepSeek Sparse Attention layers; zero full-attention layers |
| DSA selection | Top 2,048 tokens per query, GQA with 4 KV groups, 16 indexer heads |
| Training | 3.25T tokens: 50B indexer warm-up, 3T sparse-attention training, 200B LR decay |
| Precision | FP8 mixed-precision inference; ~315 GB weights |
| License | MIT (weights + inference code) |
| API price | $0.10 in / $0.40 out / $0.01 cache read per MTok |
| Inference | NaiveRT: 50 tok/s per user Standard, up to 2,000 tok/s Ultrafast |
Why the attention design is the story
Long-context MoE models from 2025-2026 typically use mostly sliding-window layers with a handful of global-attention layers to keep long-range information. At a million tokens those few global layers dominate decode cost. NaiveAI swapped them for DeepSeek Sparse Attention (DSA), which the DeepSeek team introduced with V4: a small indexer scores the whole history, then the main attention only reads the 2,048 highest-scoring tokens. NaiveAI also replaced DeepSeek’s MLA with grouped-query attention (four KV groups), which is simpler to serve on standard stacks.
The practical result is that per-token decode cost does not climb with context length the way it does in models that keep full-attention layers. The trade-off is that the indexer still scans the full history and the full KV cache is kept in memory, so memory footprint at 1M tokens is not reduced — only compute and memory bandwidth are.
How it was evaluated
NaiveAI’s model card says the coding and agentic evaluations were run inside Claude Code 2.1.207 with a 1M-token context, temperature 1.0, top-p 0.95, with the harness exposing only basic file I/O and Bash tools. Comparison scores are pulled from vendor pages for GLM-5.3, GLM-5.3-Flash, Kimi K3, Qwen 3.8 Max, Hy4-preview, DeepSeek V4.1 Flash, Step-5-preview and MiniMax M3, plus SWE-Bench Pro numbers for GPT-5.6 Sol, Opus 5 and Opus 5.5 from OpenAI’s and Anthropic’s own posts. The AI R&D suite (SOL-ExecBench, NanoChat AutoResearch, NanoGPT SpeedRun) uses NaiveAI’s in-house harness. As of September 28, 2026 there is no independent reproduction and no Artificial Analysis Intelligence Index score, so treat the charts as vendor-reported.
Where it sits in the open flash tier
The interesting comparison is not against frontier closed models but against the other 1M-context open MoE “flash” models that shipped in September 2026: Xiaomi’s MiMo-V2.6-Flash (310B/15B, $0.14/$0.28), DeepSeek V4.1 Flash (552B, $0.15/$0.60 off-peak) and Z.ai’s GLM-5.3 Flash ($0.15/$0.50, weights not shipped). Naive-N0.5-Flash undercuts all three on input price and shares a base lineage with MiMo. Full comparison: Naive-N0.5-Flash vs MiMo-V2.6-Flash vs DeepSeek V4.1 Flash vs GLM-5.3 Flash.
How to run it
pip install "transformers[torch,kernels]>=5.17.0"
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "NaiveAI/Naive-N0.5-Flash-FP8"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True,
dtype="auto", device_map="auto")
Recommended sampling is temperature 1.0 and top-p 0.95. Budget roughly 315 GB of GPU memory for weights plus KV cache; SGLang is credited in the acknowledgments, so expect that path to be the served one. For a single-GPU alternative, see best open-weight models for a single GPU 2026.
Caveats
- Vendor-reported benchmarks only, as of September 28, 2026.
trust_remote_code=Trueis required — audit the modeling code before running it on anything sensitive.- The API “will also be provided”; check availability before designing around the $0.10/$0.40 rate.
- The “AI did the research” framing is a claim about process, not a verifiable property of the weights.
Last verified: September 28, 2026.