AI agents · OpenClaw · self-hosting · automation

Quick Answer

Qwen-Audio-3.1 vs Grok Transcribe 2 vs Deepgram vs Eleven

Published:

The short answer

Alibaba’s Qwen-Audio-3.1 release on September 23, 2026 is a price event more than a model event. The five-model refresh is solid, particularly the single-model ASR covering 30 languages and 16 Chinese dialects, but the headline is the repricing: ASR down as much as 95%, Realtime down about 85%, TTS down about 70%. That lands five days after SpaceXAI’s Grok Voice Transcribe 2.0 set the Western floor at $0.10 per audio hour, and it undercuts ElevenLabs on synthesis by an order of magnitude. If your audio is Chinese or code-switches into Chinese dialects, Qwen is now the default. If it is English telephony, Grok still leads on accuracy and price clarity; if you need the deepest voice library, ElevenLabs.

Side by side

Qwen-Audio-3.1Grok Voice Transcribe 2.0Deepgram Nova-3ElevenLabs (Scribe v2 / Eleven v3)
MakerAlibaba (Tongyi Lab)SpaceXAIDeepgramElevenLabs
ReleasedSeptember 23, 2026September 18, 20262025, updated 2026Scribe v2 March 2026
ScopeASR + ASR-Next + TTS + TTS-Next + RealtimeSTT onlySTT (plus separate TTS/agent products)STT + TTS + agents
ASR price0.8 yuan in / 2.7 yuan out per M tokens (ASR-Flash)$0.10/hr batch, $0.20/hr streaming$0.26/hr batch, $0.29/hr streamingScribe v2 $0.22/hr batch, $0.39/hr realtime
TTS price1.5 yuan in / 12 yuan out per M tokens (TTS-Flash)n/an/a~$100 per M chars (v3), ~$50 (Flash)
ASR languages30 + 16 Chinese dialects, one modelMultilingual30+90+
DiarizationSpeaker labels + timestamps in outputIncludedIncludedIncluded, up to 48 speakers
Streaming latency~160 ms first characterStreaming, first on AA leaderboardSub-300 ms~150 ms
Open weightsNo (hosted on Model Studio)NoNoNo
Public benchmark4.55% avg CER on KeSpeech + WSYue dialect sets#1 of 32 on AA streaming WERVendor-reported~2.3% AA-WER (Scribe v2)

What Alibaba shipped on September 23, 2026

Five models, one stack:

  • Qwen-Audio-3.1-ASR: end-to-end recognition that emits speaker labels, timestamps and text together, handles overlapping speech and short interjections, keeps names and abbreviations consistent across long audio, and includes a post-transcription polishing pass that strips filler words. 30 languages and 16 Chinese dialects in a single model; 10.38% average CER on Alibaba’s internal 16-dialect set, with Wenzhou and Suzhou dialects called out as far ahead of competitors.
  • Qwen-Audio-3.1-ASR-Next: audio understanding beyond words, including emotion, music, environmental sound, event localization and speaker identification. Its API was not yet live at launch.
  • Qwen-Audio-3.1-TTS: 16 languages (seven new) plus 20 Chinese dialect regions, cross-language timbre transfer so one voice carries across Mandarin, Cantonese, English and Japanese, instruction-based control of emotion and pace, and tolerance for noisy reference audio.
  • Qwen-Audio-3.1-TTS-Next: the “Audiogen” tier. Voice, sound effects and background ambience generated together from text, timestamps and reference audio; multi-speaker dialogue and podcasts; 48 kHz; up to 3,000 input characters and 240 seconds of output for podcasts (120 otherwise); 3 requests per second.
  • Qwen-Audio-3.1-Realtime: full-duplex conversation, being wired into Alibaba’s own hardware (QwenNote A2, the Eva desktop robot, Qwen AI glasses) and adjusting pace and tone when it detects a low mood.

Alibaba also open-sourced Qwen-Audio-Agent, a framework for real-time voice agents, and showed Qwen3.8-LiveTranslate with latency under 2.5 seconds. For the omnimodal LLM side of the same family, see Qwen3.8-Omni-Flash vs Gemini 3.8 Flash.

The pricing problem: tokens versus hours

Every Western vendor in this comparison bills speech-to-text per audio hour. Alibaba bills per million audio tokens, and its audio tokenizer rate is model-specific. That makes the “95% cut” real but not directly comparable: Qwen-Audio-3.1-ASR-Flash at 0.8 yuan input and 2.7 yuan output per million tokens (about $0.11 and $0.38 at late-September 2026 exchange rates) is cheap in absolute terms, but the per-hour figure depends on how many tokens an hour of your audio becomes. Run a one-hour sample through Model Studio and read the token count off the bill before you tell procurement it beats Grok’s $0.10.

For TTS the comparison is easier because ElevenLabs bills per character and Qwen’s TTS input is text tokens: TTS-Flash at 1.5 yuan in / 12 yuan out per million tokens is roughly one to two orders of magnitude below Eleven v3’s ~$100 per million characters. The international listing for TTS-Next ($0.848 in / $1.696 out per million tokens) confirms the scale.

Where each wins

Qwen-Audio-3.1 wins on Chinese, on price for synthesis, and on being one stack: ASR, understanding, TTS, sound design and full-duplex from one vendor with one bill. The KeSpeech/WSYue result (4.55% CER, ahead of Doubao-ASR and Tencent Hy-ASR-3.0 on 6 of 11 subsets) is the number to cite. The caveats: hosted only, ASR-Next API not yet available, and token pricing that needs conversion.

Grok Voice Transcribe 2.0 wins on English and multilingual streaming accuracy and on price legibility. $0.10 per hour batch, $0.20 streaming, diarization and word timestamps included, first of 32 models on the Artificial Analysis streaming leaderboard as of September 18, 2026. It is STT only. Full breakdown in Grok Transcribe 2.0 vs Deepgram vs AssemblyAI vs Scribe v2.

Deepgram Nova-3 wins on maturity: custom-model training, keyword boosting, on-prem deployment and the largest installed base of voice-agent stacks. It is now the most expensive STT of the four at $0.26 per hour batch.

ElevenLabs wins on voice: the largest shared voice library, instant cloning from about 30 seconds of audio, 74 TTS languages and the top Provider Voice Arena Elo (1,168 as of August 2026). Scribe v2 is a credible STT at $0.22 per hour, but the reason to pay ElevenLabs is Eleven v3, and that is exactly the product Qwen just repriced against.

Decision rule

  • Chinese or dialect-heavy audio, or mixed Mandarin/Cantonese/English: Qwen-Audio-3.1, no contest.
  • English telephony and meetings, lowest clear per-hour price: Grok Voice Transcribe 2.0.
  • Existing Deepgram agent with custom models: stay unless the audio bill is material.
  • Synthesis at volume where the invoice decides: trial Qwen-Audio-3.1-TTS against Eleven v3 Flash on your own scripts; the gap is large enough to justify a migration test.
  • Voice library, cloning and Western-language breadth: ElevenLabs.

Last verified: September 25, 2026.

Sources