Best Text-to-Speech APIs 2026: Ranked by Quality and Price
The ranking (September 2026)
Rankings weigh independent blind-preference data (Artificial Analysis Speech Arena), documented capabilities (latency, streaming, cloning, languages, controls) and price. Vendor latency claims are model latency, not end-to-end round trips — benchmark your own.
| # | API / model | Best for | Arena signal | Latency (vendor) | Price (list) |
|---|---|---|---|---|---|
| 1 | Cartesia Sonic 3.6 (beta; 3.5 stable) | Real-time voice agents | #1 Provider Voice 1,283 Elo; #1 Controlled Voice 1,123 | Sub-90 ms TTFA | ~$49 / 1M chars normalised; plans Free, Pro $5, Startup $49, Scale $299 (~10,667 min) |
| 2 | ElevenLabs — Eleven v3 / Multilingual v2 / Flash v2.5 | Creator ecosystem, cloning, dubbing, dialogue | Eleven v3 #3 on Controlled Voice | Flash v2.5 ~75 ms | ~$100 / 1M chars normalised; Starter $5 → Scale $330/mo |
| 3 | Google Gemini 3.1 Flash TTS (Preview) | Prompt-directed multi-speaker audio | ~1,211 Elo | Streaming | $1 / 1M text in + $20 / 1M audio out (~$1.80 per finished hour) |
| 4 | Qwen Audio 3.0 TTS Plus | Prerecorded narration, ads, character lines | Top provider-voice signal (narrow margin) | Lower throughput; 3 req/s limit | Alibaba Cloud Model Studio, per-character |
| 5 | Inworld Realtime TTS-2 | Multilingual real-time apps with lip-sync | — | ~200 ms median | Usage tiers; pin model IDs (TTS-2 vs 1.5 Max/Mini) |
| 6 | OpenAI gpt-4o-mini-tts | Teams already on OpenAI | — | Streaming | $0.60 / 1M text tokens + $12 / 1M audio tokens (~$0.015/min) |
| 7 | Deepgram Aura-2 | Voice-agent stacks already on Deepgram STT | — | Low | ~$30 / 1M chars |
| 8 | Google Chirp 3 HD | GCP-native, many locales | — | — | $30 / 1M chars after 1M free/month |
| 9 | MiniMax Speech 2.8 HD / Turbo | Long-form, 40 languages, cloning | — | Streaming | $100 / 1M chars HD; $60 Turbo |
| 10 | Fish Audio S2.1 Pro Free | Free prototyping | — | No TTFA guarantee on free route | $0 (free route); paid route for production |
1. Cartesia Sonic 3.6 — best for real-time agents
Cartesia released Sonic 3.6 on August 18, 2026, about three months after Sonic 3.5. It now holds #1 on both Artificial Analysis speech leaderboards: 1,283 Elo on Provider Voice and 1,123 on Controlled Voice. The second result is the one that matters — that board clones every model onto the same eight reference voices, isolating the synthesis engine from the voice catalogue. Sonic 3.5 is second and ElevenLabs Eleven v3 third.
Sonic runs on state space models rather than transformers; Cartesia states sub-90 ms time-to-first-audio. Production features are built for agent transcripts: inline expression tags ([laughter]), instant cloning from ~10 seconds of audio, custom pronunciation dictionaries with IPA overrides, speed/volume/emotion parameters, native alphanumerics for order and phone numbers, and a WebSocket flow that ingests LLM text fragments while preserving context. Sonic 3.5 supports 42 languages.
Caveats: 3.6 is beta on Cartesia’s hosted API only (docs still list 3.5 as stable and partners like LiveKit carry 3.5); there are no open weights. Cartesia sells credits, not characters — Artificial Analysis normalises it at ~$49 per 1M characters, half of Eleven v3; Scale at $299/month includes ~10,667 TTS minutes and 15 concurrent requests; Line voice agents bill separately at $0.06/min. Commercial use starts at the $5 Pro tier.
2. ElevenLabs — best creator ecosystem
ElevenLabs is no longer the top voice in every blind comparison, but it has the most complete product around its API. Three models cover different jobs — and picking the wrong one is the most common mistake:
- Eleven v3 — expressive speech, dialogue and 70+ languages; use for narration and character work.
- Multilingual v2 — stable long-form across 29 languages; use for audiobooks and consistent output.
- Flash v2.5 — ~75 ms model latency, 32 languages, larger character limit; use for agents.
Around them: a voice library, instant and professional cloning, voice design, dubbing, pronunciation control and Text to Dialogue. Plans run Free (10K chars/month, non-commercial), Starter $5, Creator $22, Pro $99, Scale $330 per month, with Artificial Analysis normalising Eleven v3 at ~$100 per 1M characters — the premium you pay for the ecosystem.
3. Gemini 3.1 Flash TTS — best controllable multi-speaker
Google’s Gemini TTS generates single- or multi-speaker audio from natural-language direction (“accent, style, pace, tone”) plus expressive audio tags, which makes it the most interesting option for podcasts, lessons and scripted two-person dialogue. Its Speech Arena Elo is ~1,211. Pricing is token-based: $1 per 1M text input tokens and $20 per 1M audio output tokens; at Google’s documented 25 audio tokens per second that is about $1.80 per finished audio hour for output. The word to respect is Preview: Google warns preview models can change and carry tighter rate limits.
4. Qwen Audio 3.0 TTS Plus — best raw voice quality signal
Alibaba’s Qwen Audio 3.0 TTS Plus sits at the top of Artificial Analysis’ provider-voice leaderboard by a narrow margin. It takes natural-language instructions for tone, speed, emotion and timbre, inline tags such as [sad] or [excited], non-verbal effects, cloning and real-time synthesis. The costs are throughput (lower than real-time-focused rivals), a documented 3 requests/second job-submission limit, and separate international/China deployments. Pick it for prerecorded assets, not a phone agent.
5. Inworld Realtime TTS-2 — best multilingual real-time
Inworld lists 200+ languages and locales, instant cloning, natural-language steering and timestamp output with phonemes and visemes — useful for avatars, lip-sync and language learning — at ~200 ms median latency. The product family also includes TTS 1.5 Max (stability) and 1.5 Mini (latency); the old inworld-tts-1 IDs were retired in June 2026 and now route to 1.5. Pin the model ID deliberately.
6–10. The value and niche tiers
- OpenAI gpt-4o-mini-tts — the current OpenAI TTS model, priced per token ($0.60/1M text in, $12/1M audio out ≈ $0.015 per minute), with streaming and instructable style; preset voices are English-optimised and there is no native multi-speaker mode. Right if you are already on the OpenAI platform.
- Deepgram Aura-2 — ~$30 per 1M characters, low latency, and it sits next to the Deepgram STT most voice-agent builders already run, removing a vendor from the diagram.
- Google Chirp 3 HD — $30 per 1M characters after a 1M-character monthly free allowance (Neural2 at $16); the GCP-native choice with broad locale coverage.
- MiniMax Speech 2.8 HD / Turbo — 40 languages, seven named emotions, interjection tags, 10,000-character requests; $100 (HD) / $60 (Turbo) per 1M characters. Speech 2.6 and Speech 02 are now legacy.
- Fish Audio S2.1 Pro Free — same endpoint and model as s2.1-pro at $0 for prototyping, minus TTFA and data-processing guarantees; 64+ emotion/style cues, HTTP and WebSocket.
How to choose
- Voice agent where a 300 ms pause feels like a bug: Cartesia Sonic 3.6/3.5, then ElevenLabs Flash v2.5.
- Creator product needing cloning, dubbing and editing tools: ElevenLabs.
- Two-speaker scripted content: Gemini 3.1 Flash TTS (accept preview status).
- Highest-fidelity prerecorded narration: Qwen Audio 3.0 TTS Plus or MiniMax Speech 2.8 HD.
- Cheapest acceptable quality at volume: Speechify Simba, OpenAI gpt-4o-mini-tts, Deepgram Aura-2, Chirp 3 HD.
- Prototype for free: Fish Audio s2.1-pro-free, ElevenLabs Free tier, Cartesia Free.
Whatever you pick, compare cost per finished audio minute, not per-character list price — token-billed models (Gemini, OpenAI) and credit-billed ones (Cartesia, ElevenLabs) do not convert cleanly.