AI agents · OpenClaw · self-hosting · automation

Quick Answer

Best Local Speech-to-Text Models 2026: Top 7 Ranked

Published:

The short answer

As of October 2026 the local speech-to-text field has split by workload: Voxtral Realtime for multilingual streaming, Parakeet Ultra for Apple-silicon dictation, Canary-Qwen 2.5B for English accuracy, Whisper large-v3-turbo for language breadth, and Whistle or Moonshine for devices measured in megabytes. Whisper is no longer the accuracy leader in any category, but it is still the only model on this list that covers 99 languages. Facts verified October 4, 2026.

The ranking

#ModelMaker / dateParams / sizeLanguagesStreamingLicenceBest for
1Voxtral RealtimeMistral, Feb 4, 20264B (~2.5 GB Q4)13✅ native, 80 ms–2.4 sApache 2.0Multilingual live transcription, 3-hour audio
2Parakeet UltraMoondream, Sep 22, 2026600M25 EuropeanChunked (VAD head)NVIDIA/Moondream openApple-silicon and GPU dictation
3Canary-Qwen 2.5BNVIDIA, 20252.5BEnglish❌CC-BY-4.0Maximum English accuracy
4Whisper large-v3-turboOpenAI, Oct 2024809M (~3 GB Q5)99❌ (chunked)MITLanguage breadth, ecosystem
5Parakeet Unified 0.6BNVIDIA, Apr 2026600MEnglish✅ ~160 msNVIDIA openEnglish streaming with punctuation
6Moonshine Tiny v2Useful Sensors, 202541.9 MBEnglish✅MITRaspberry Pi-class edge
7WhistleCactus, Oct 2, 202616.9 MB7❌ (30 s pass)OpenMicrocontrollers, wearables, voice-to-tool-calls

Why each one is ranked where it is

1. Voxtral Realtime (Mistral). Launched February 4, 2026 as half of Voxtral Transcribe 2 (the other half, Mini Transcribe V2, is API-only). It is the only open model here with a causal audio encoder trained for streaming, so latency is a dial from 80 ms to 2.4 s rather than a chunking hack, and it handles about three hours of audio natively through a 131K-token window. On FLEURS across its 13 languages it reports roughly 5.9% average WER against Whisper large-v3’s 7.4%. The costs: 16 GB VRAM in BF16 (about 2.5 GB quantised), and 13 languages is a hard ceiling.

2. Parakeet Ultra (Moondream). Moondream released Ultra and Redux on September 22, 2026 as post-trained variants of NVIDIA’s Parakeet TDT 0.6B V3, keeping V3’s architecture and tokenizer and adding a small voice-activity head so the Photon runtime can split long recordings at pauses. Ultra is tuned for accuracy, Redux compressed for size; both cover 25 European languages with automatic detection, punctuation, capitalisation and word timestamps. On Moondream’s own tests Ultra beats V3 on English, FLEURS, business speech (AMI, VoxPopuli, Earnings-22) and noisy audio. For Mac and iPhone dictation apps this is now the default pick; for a CPU-only server, Redux.

3. Canary-Qwen 2.5B (NVIDIA). Still the open-model leader on the English Open ASR Leaderboard at roughly 5.1–5.6% average WER depending on snapshot. It is English-only and batch-only, so it is the model you reach for when transcribing recorded English where every word matters — depositions, interviews, podcasts — and not for live use.

4. Whisper large-v3-turbo (OpenAI). Four decoder layers instead of 32 makes it about 8× faster than large-v3 at nearly the same accuracy, and whisper.cpp runs it on an 8 GB laptop or an iPhone. It has no native streaming and processes audio in 30-second windows, so long-form needs WhisperX or whisper.cpp’s VAD. It stays at number four because for anything outside the 13 Voxtral or 25 Parakeet languages it is the only serious option, and because three years of tooling (faster-whisper, WhisperX, MLX ports) is worth a lot in production.

5. Parakeet Unified 0.6B (NVIDIA). April 2026, English-only RNN-T model that does both whole-recording and streaming from one checkpoint. NVIDIA’s card reports 5.91% WER whole-recording and 6.29% streaming at 1.12 s latency, with latency configurable down to about 160 ms. The companion Parakeet Realtime EOU 120M streams at 80–160 ms and emits end-of-utterance markers for voice agents, at 9.30% WER and no punctuation.

6. Moonshine Tiny v2 (Useful Sensors). Built for streaming on constrained hardware; latency scales with audio length instead of padding to 30 seconds like Whisper, which is why it beats Whisper on time-to-first-token on short clips. English only. The right choice for a Raspberry Pi or a browser extension that needs live captions.

7. Whistle (Cactus Compute). The newest entry at 16.9 MB, a single file with no dependencies, seven languages, word timestamps with probabilities, keyword biasing, and 11 ms to first token on an M4 Pro. It beats the 145 MB Whisper Base on LibriSpeech, SPGISpeech, Earnings-22 and FLEURS while losing on TED-LIUM, AMI and MLS. Its unique trick is sharing a C++ engine with Cactus’s Needle tool-calling model so one binary turns a clip into function calls. Last on the list only because 30-second passes and seven languages limit the use cases; first on the list if your device has 20 MB to spare. Full comparison: Whistle vs Whisper Base vs Moonshine.

How to choose in 30 seconds

  • Need a language outside Voxtral’s 13 or Parakeet’s 25: Whisper large-v3-turbo. Nothing else covers it.
  • Live captions, multilingual, GPU available: Voxtral Realtime.
  • Dictation app on Mac/iPhone: Parakeet Ultra (or Redux for the smallest download).
  • Batch English transcription, accuracy above all: Canary-Qwen 2.5B.
  • Voice agent needing end-of-turn detection: Parakeet Realtime EOU 120M, then a bigger model for the final transcript.
  • Edge device under 50 MB: Whistle (multilingual, tool calls) or Moonshine Tiny (English streaming).

Read every WER with its benchmark attached. Vendor numbers on this page come from different test sets, normalisers and snapshots; a 0.5-point gap between two vendors’ self-reported averages is noise. The Open ASR Leaderboard is the one place where the same harness scores every model.

Related: best speech-to-text APIs 2026 ranked for the hosted options, Whisper vs MAI-Transcribe-1 vs Deepgram, best on-device AI models 2026.

Last verified: October 4, 2026. WER figures are the vendors’ published numbers on the datasets named; the Open ASR Leaderboard is the cross-vendor reference.

Sources