MAI-Transcribe-2-Streaming vs Grok Transcribe 2 vs Deepgram
The short answer
Microsoft’s MAI-Transcribe-2-Streaming, released October 1, 2026, is the new accuracy leader for real-time speech-to-text — #1 on Artificial Analysis for both final and partial transcripts, with first partials in about 100 ms — but at $0.54 per hour it costs 2.7x Grok Voice Transcribe 2.0’s $0.20. Pick Microsoft for latency-sensitive agents, Grok for cost, Deepgram for a mature production stack, Scribe v2 for language breadth. Facts verified October 2, 2026.
Side by side
| MAI-Transcribe-2-Streaming | Grok Voice Transcribe 2.0 | Deepgram Nova-3 | ElevenLabs Scribe v2 Realtime | |
|---|---|---|---|---|
| Maker | Microsoft AI | SpaceXAI (formerly xAI) | Deepgram | ElevenLabs |
| Released | Oct 1, 2026 | Sep 18, 2026 | 2025, updated 2026 | 2026 |
| Streaming price / hour | $0.54 (intro to Dec 31, 2026) | $0.20 | $0.29 promo / $0.46 regular | $0.39 ($0.28 annual Business) |
| Batch price / hour | Streaming-only model | $0.10 | $0.26 | $0.22 (Scribe v2 batch) |
| Languages | 60, continuous auto-detect | Dozens, mid-recording switching | Multilingual | 90+ |
| First partial | ~100 ms | Not published | Sub-300 ms (vendor) | ~150 ms |
| Diarization | Not stated at launch | Included | Included | Included, up to 48 speakers |
| AA streaming rank | #1 accuracy (final + partial) | #1 at its Sep 18 launch; now displaced | Mid-pack | ~2.3% AA-WER (non-streaming index) |
| Open weights | No | No | No | No |
| Paired TTS | MAI-Voice-2.1 / 2.1-Flash | Grok Voice | Aura-2 | Eleven v4 / v4 Turbo |
| Best for | Lowest-latency agents, live captions | Cheapest production audio | Incumbent stacks, on-prem | Multilingual, one vendor for STT+TTS |
Prices are pay-as-you-go list rates read from vendor pages on October 2, 2026; volume and annual tiers are lower.
MAI-Transcribe-2-Streaming — what Microsoft shipped
Microsoft AI released MAI-Transcribe-2-Streaming on October 1, 2026, alongside two text-to-speech models (MAI-Voice-2.1 and MAI-Voice-2.1-Flash). Four claims matter:
- #1 on Artificial Analysis for streaming accuracy, on both final transcripts and partial (in-progress) hypotheses, and on the Pareto frontier of AA’s accuracy-versus-latency plot — meaning no listed model is both more accurate and faster.
- ~100 ms to first partial. The model emits a hypothesis just over 100 ms after receiving audio, revises as context arrives, and commits a stable transcript immediately. Microsoft’s internal tests say words appear 2x faster than its closest competitor for dictation and subtitling.
- 60 languages with continuous automatic detection, so a caller can switch languages mid-sentence without a config change.
- $0.54 per hour of audio, introductory through the end of 2026. Microsoft has not published the post-promo rate; budget higher.
The design target is explicit: a voice agent is a hear-understand-decide-speak loop, and partials let the agent start reasoning or calling tools before the speaker finishes. Microsoft’s Chatter demo in the MAI Playground pairs it with MAI-Voice-2.1-Flash (150 ms end-to-end for 45 s of audio, $15 per 1M characters). This is Microsoft’s second-generation in-house STT after MAI-Transcribe-1, and part of the MAI self-sufficiency push away from OpenAI models.
Grok Voice Transcribe 2.0 — the price floor
SpaceXAI’s Grok Voice Transcribe 2.0 (September 18, 2026) held the #1 streaming-accuracy slot for 13 days. It remains the cheapest serious option: $0.20/hr streaming, $0.10/hr batch, with diarization, word timestamps with confidence, up to 100 key terms, 8-channel audio and smart turn detection included rather than billed as add-ons. SpaceXAI’s internal sets showed telephony first-final WER at 2.7% and short-phrase multilingual WER at 6.8%. Endpoint api.x.ai/v1/stt; API-only, no open weights. Full breakdown in Grok Transcribe 2.0 vs Deepgram vs AssemblyAI vs Scribe v2.
Deepgram Nova-3 — the incumbent
Most 2025-era voice-agent stacks were built on Nova-3. Streaming is $0.0048/min (~$0.29/hr) at a promotional rate against a regular $0.0077/min ($0.46/hr) with no published end date — model your budget at the regular rate. Batch is $0.0043/min ($0.26/hr); new accounts get $200 credit. Deepgram’s advantages are maturity, keyword boosting, custom vocabularies and on-prem deployment for enterprise. Its disadvantage is that two newer models now beat it on both accuracy and price.
ElevenLabs Scribe v2 Realtime — breadth
Scribe v2 Realtime is $0.39/hr ($0.28 on annual Business plans) with ~150 ms latency, 90+ languages, diarization for up to 48 speakers and non-speech event tags; keyterm prompting (+$0.05/hr) and entity detection (+$0.07/hr) are add-ons. Its case is consolidation: ElevenLabs shipped Eleven v4 and v4 Turbo TTS on September 28, 2026, so one vendor covers both ends of a voice agent. See MAI-Voice-2.1 vs Eleven v4 vs Cartesia Sonic 3.6 for the TTS side.
Also in the market
- AssemblyAI Universal-3.5 Pro — $0.45/hr realtime, $0.21/hr async, billed per second; diarization +$0.12/hr realtime; strong on entity accuracy and code-switching across 18 languages; Voice Agent API bundles STT+LLM+TTS at $4.50/hr.
- Meta Muse Voice Transcribe — $0.18/hr streaming and batch (September 1, 2026), zero-data-retention option at the same price.
- OpenAI GPT Transcribe — $0.0045/min ($0.27/hr); legacy Whisper API $0.006/min ($0.36/hr). Whisper large-v3 remains the self-hosting default because none of the models above have open weights.
How to choose
- Agent where a 300 ms pause feels like a bug: MAI-Transcribe-2-Streaming, paired with MAI-Voice-2.1-Flash or Eleven v4 Turbo.
- High-volume call transcription where cost dominates: Grok Voice Transcribe 2.0 — 2.7x cheaper than Microsoft and diarization is free.
- Existing Deepgram integration that works: stay on Nova-3, but renegotiate — your vendor is now mid-pack on price and accuracy.
- Need 90+ languages or one vendor for STT and TTS: ElevenLabs Scribe v2.
- Audio cannot leave your servers: self-hosted Whisper large-v3; none of these four qualify.
A warning on leaderboards: Grok held #1 for 13 days. Artificial Analysis rankings reorder with each release; the durable facts are latency architecture and price, both in the table above.
Last verified: October 2, 2026. Microsoft figures from the October 1 announcement; Grok, Deepgram, ElevenLabs and AssemblyAI prices from vendor pricing pages.