Whistle vs Whisper Base vs Moonshine Tiny: Tiny STT (2026)
The short answer
Cactus Whistle (October 2, 2026) is the smallest credible on-device speech-to-text model available: 16.9 MB, 11.1 ms to first token, and ahead of the 145.3 MB Whisper Base on LibriSpeech, SPGISpeech, Earnings-22 and FLEURS. Whisper Base still wins on TED-LIUM, AMI and MLS and covers 99 languages to Whistle’s seven. Moonshine Tiny v2 sits between them at 41.9 MB and is the only one of the three built for genuine streaming. Facts verified October 4, 2026 from the Cactus release post and the models’ published cards.
Side by side
| Whistle | Whisper Base | Moonshine Tiny v2 | |
|---|---|---|---|
| Maker | Cactus Compute | OpenAI | Useful Sensors |
| Released | Oct 2, 2026 | Sep 2022 (base) | 2025 (v2) |
| Size on disk | 16.9 MB | 145.3 MB | 41.9 MB |
| Languages | 7 (en, de, fr, es, it, nl, pl) | 99 | English only |
| Max clip per pass | 30 s | 30 s (padded) | Streaming |
| Time to first token (10 s clip, M4 Pro CPU) | 11.1 ms | 73.2 ms | 22.8 ms |
| Decode speed | 1,319 tok/s | 266 tok/s | 262 tok/s |
| Word timestamps | ✅ with probability | Via tooling | ✅ |
| Keyword biasing | ✅ (Aho-Corasick) | ❌ | ❌ |
| Dependencies | None (single C++ engine) | PyTorch / whisper.cpp | ONNX / own runtime |
| Licence | Open (see HF card) | MIT | MIT |
Speed figures are Cactus’s, measured on an Apple M4 Pro CPU with each model on its official runtime at defaults: Whistle’s engine at 5 beams, openai-whisper, and moonshine-voice non-streaming over whole audio.
What makes Whistle 16.9 MB
Whistle is not a distilled Whisper. It reuses the architecture of Cactus’s Needle tool-calling model: eight Simple Attention encoder blocks (non-causal, so a frame at 3 s can attend to a frame at 12 s) and eight laddered decoder blocks at width 512 with 8 query heads to 2 KV heads. The only speech-specific addition is a gated cross-attention per decoder layer that reads the encoder output. Those cross-attention projections run once when the clip arrives and are then held for the whole decode, so five-beam search costs five short transcript caches rather than five passes over the audio.
Two engineering choices matter for embedded use:
- The decoder is a ladder. Every depth from 2 layers up was trained as its own model, and
--audio-depthpicks one at load time. You trade accuracy for latency without downloading a different file. The encoder always runs all eight blocks. - Silence short-circuits. The engine measures loudness range before decoding and returns an empty transcript below threshold, never entering beam search. On a wake-word device that is most of the calls.
The vocabulary is 8,192 text pieces plus seven language tokens, so the detected language comes back as a token rather than a side channel. Transcripts are capped at 320 tokens.
Accuracy, honestly
Whistle’s WER is measured by Cactus over 86,174 utterances with Whisper normalizers; Whisper’s and Moonshine’s are the figures their authors published for multilingual checkpoints. Cactus verified by audio checksum and speaker ID that no test audio leaked into training. With that caveat:
- Whistle ahead: LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22, FLEURS average.
- Whisper Base ahead: TED-LIUM, AMI, MLS average. (Whisper’s AMI number is AMI-IHM, a different subset.)
- Moonshine: English only, so it has no FLEURS or MLS numbers to compare.
The pattern is consistent with the design: Whistle is strong on read and business speech and on the seven languages it targets, weaker on long-form lecture audio where Whisper’s bigger decoder helps. If your corpus looks like TED talks, Whisper Base’s extra 128 MB buys real accuracy. If it looks like voice commands and earnings calls, it does not.
Speech straight into tool calls
The feature that separates Whistle from every other tiny ASR model is the engine, not the weights. needle_load reads whichever model a .cact file holds, so:
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav
returns one JSON object with both the function calls and the speech fields (audio_text, audio_language) — the caller never touches a transcript. In Python it is pip install cactus-needle and needle.transcribe("clip.wav")["text"]; keywords=["Siobhan", "Krzysztof"] biases the beam search toward names your users actually say. For a smart-home or automotive voice stack that is the difference between two models with glue code and one binary.
Which to pick
- Microcontroller, wearable or any device where 16.9 MB vs 145 MB is the decision: Whistle.
- Voice commands feeding a local tool-calling agent: Whistle + Needle, one runtime.
- English-only live captions with true streaming: Moonshine Tiny v2. Whistle and Whisper both work in 30-second passes.
- Any language outside Whistle’s seven, or lecture/long-form audio: Whisper Base via whisper.cpp — or step up to a larger local model if the device allows.
Related: best local speech-to-text models 2026, Whisper vs MAI-Transcribe-1 vs Deepgram, best on-device AI models 2026.
Last verified: October 4, 2026. Benchmark and latency figures are Cactus Compute’s published results; Whisper and Moonshine WER figures are from their authors.