Updated September 23, 2026. Three things changed since our July edition. Cartesia Sonic 3.6 (August 27) took first place on the independent Artificial Analysis TTS arena. OpenAI GPT-Live-1 shipped as an API at a flat $0.05 per minute. Google Gemini 3.8 Live (September 15) now leads the speech-to-speech rankings. ElevenLabs is no longer the quality leader on blind votes, but it is still the most complete voiceover product.
This guide is for founders and marketers who need voiceovers, podcast intros, ads, and explainer narration, plus teams building a voice agent. Picks come first, then the ranking, prices, and how to choose.
Need a voiceover today, not a vendor evaluation? Paste your script into Teamday Audio and get a downloadable voiceover from Gemini TTS, kept next to the video and files it belongs to.
Create your first voiceover →Quick Answer: The Best AI Voice Model for Each Job
| Your job | Start with | Why |
|---|---|---|
| Ad, explainer, or brand voiceover | ElevenLabs v3 | Expressive delivery, 70+ languages, voice cloning, mature editor; $0.10 per 1K characters |
| Highest blind-test quality via API | Cartesia Sonic 3.6 | #1 on the Artificial Analysis TTS arena (ELO 1273); under 90 ms latency claim |
| Best quality per dollar | Inworld Realtime TTS-2 or Qwen-Audio-3.0-TTS | #3 and #2 on the arena at a fraction of ElevenLabs v3's list price |
| Cheap high-volume narration | Fish Audio S2.1 Pro or Deepgram Aura-2 | $15 per 1M bytes (free tier until Nov 30, 2026) / $0.030 per 1K characters |
| Live voice agent on the phone or web | OpenAI GPT-Live-1 or Gemini 3.8 Live | Top of the speech-to-speech index; GPT-Live-1 is $0.05/min plus the reasoning model |
| Voiceovers inside Google Cloud | Gemini 3.1 Flash TTS (preview) | #8 on the arena, $1 text in / $20 audio out per 1M tokens |
| Self-hosted, commercial-safe | Qwen3-TTS or Kokoro-82M | Apache 2.0 weights; Qwen3-TTS clones from 3 seconds of audio |
Do not choose from a 20-second demo. Test your own script, your product names, and every target language.
September 2026 Ranking: Text-to-Speech Quality
Artificial Analysis runs blind A/B listening tests and converts the votes into ELO scores. It is the best public quality signal, but it measures short English clips, not your brand voice, pronunciation of your product names, or long-form consistency.
| Rank | Model | Vendor | Arena ELO | Released | Vendor list price |
|---|---|---|---|---|---|
| 1 | Sonic 3.6 | Cartesia | 1273 | Aug 2026 | Plans from $5/mo (100K credits) |
| 2 | Qwen-Audio-3.0-TTS-Plus | Alibaba | 1259 | Jul 2026 | API; AA lists $27.60 / 1M chars |
| 3 | Realtime TTS-2 | Inworld | 1245 | Aug 2026 | $25 / 1M chars on demand |
| 4 | Simba 3.2 | Speechify | 1237 | Jul 2026 | AA lists $6.60 / 1M chars |
| 5 | Luna TTS | VUI Labs | 1230 | Jun 2026 | AA lists $15 / 1M chars |
| 6 | Realtime TTS-2 Flash | Inworld | 1210 | Aug 2026 | AA lists $10.40 / 1M chars |
| 8 | Gemini 3.1 Flash TTS | 1199 | Apr 2026 | $1 text in / $20 audio out per 1M tokens | |
| 10 | v3 Conversational | ElevenLabs | 1196 | Aug 2026 | $0.05 / 1K chars |
| 11 | Sonic 3.5 | Cartesia | 1185 | May 2026 | Same plans as 3.6 |
| 14 | Speech 2.8 HD | MiniMax | 1168 | Feb 2026 | $100 / 1M chars |
| 15 | Eleven v3 | ElevenLabs | 1167 | Feb 2026 | $0.10 / 1K chars |
Source: Artificial Analysis TTS leaderboard, checked September 23, 2026. Ranks 7, 9, 12, and 13 are BreezeBlue Breeze TTS 2, StepFun StepAudio 2.5, Smallest.ai Lightning V3.1 Pro, and Soniox TTS v2. Vendor prices are from each vendor's pricing page. Where a vendor does not publish a per-character rate, we show Artificial Analysis's normalized price and label it "AA lists".
Realtime Voice Agents: Speech-to-Speech Ranking
| Rank | Model | Index score | Price |
|---|---|---|---|
| 1 | Gemini 3.8 Live Extended Thinking (High) | 82.6 | $0.005/min audio in, $0.018/min audio out |
| 2 | OpenAI GPT-Live-1 | 81.5 | $0.05/min, billed per second, plus the backend reasoning model |
| 3 | xAI Grok Voice Think Fast 2.0 | 81.3 | $0.08/min |
Source: Artificial Analysis speech-to-speech, OpenAI GPT-Live-1 model page, Gemini API pricing, xAI pricing, all checked September 23, 2026.
GPT-Live-1 is a new pricing model. It is a full-duplex voice layer that listens and speaks at the same time and hands reasoning and tool calls to a backend model you pick. It runs only on the new v1/live/sessions endpoint. The older GPT-Realtime 2.1 (July 2026) is still offered at $32 per 1M audio input tokens and $64 per 1M audio output tokens. A cheaper mini version costs $10 and $20.
Price Per Finished Minute
Most TTS vendors price per character. A spoken minute at a normal pace of about 150 words is roughly 900 characters. The table below converts list prices on that assumption. It is an estimate, not a vendor quote.
| Model | List price | ≈ Per spoken minute |
|---|---|---|
| ElevenLabs v3 / Multilingual v2 | $0.10 / 1K chars | ~$0.09 |
| MiniMax Speech 2.8 HD | $100 / 1M chars | ~$0.09 |
| ElevenLabs Flash v2.5 / v3 Conversational | $0.05 / 1K chars | ~$0.045 |
| Deepgram Aura-2 | $0.030 / 1K chars | ~$0.027 |
| Inworld Realtime TTS-2 | $25 / 1M chars | ~$0.023 |
| Mistral Voxtral TTS (API) | $0.016 / 1K chars | ~$0.014 |
| Fish Audio S2.1 Pro | $15 / 1M UTF-8 bytes | ~$0.014 (English) |
| OpenAI tts-1 | $15 / 1M chars | ~$0.014 |
| Gemini TTS, gpt-4o-mini-tts | Priced per audio token | Measure with a real script |
The generation fee is rarely what drives the cost. Retakes, pronunciation fixes, and review time usually cost more. Track total spend divided by approved finished minutes, not the list price.
The Models, One by One
Cartesia Sonic 3.6: best blind-test quality
Sonic 3.6 (August 27, 2026) is #1 on the TTS arena. Cartesia claims under 90 ms latency and covers 44 languages (61 locales). Instant voice cloning starts on the Pro plan and professional cloning on Startup. Pricing is credit-based: Pro is $5/month for 100K credits, Startup $49 for 1.25M, and Scale $299 for 8M (pricing). Built for streaming agents, it is also the strongest pure-quality pick for narration if your team is comfortable working through an API.
ElevenLabs v3: best complete voiceover product
ElevenLabs lists four TTS tiers:
- v3: expressive, 70+ languages, $0.10 per 1K characters.
- v3 Conversational: a low-latency v3 for agents, $0.05.
- Multilingual v2: consistent long-form narration, $0.10.
- Flash v2.5: about 75 ms, 32 languages, $0.05.
It no longer wins blind votes. It remains the easiest end-to-end product for a marketer: a voice library, cloning, an editor, dubbing, and music under one account.
ElevenLabs Flash v2.5, stock voice "Sarah". Generated through the API on July 13, 2026.
Alibaba Qwen-Audio-3.0-TTS and Inworld TTS-2: best quality per dollar
Qwen-Audio-3.0-TTS (July 21, 2026) comes in Plus and Flash versions. It covers 16 languages, supports voice cloning, and Alibaba claims about 300 ms first-packet latency for Flash. It is API-only, not open weights. Inworld's Realtime TTS-2 lists $25 per 1M characters on demand, dropping to $5 on Enterprise (pricing), and ranks #3. Both beat ElevenLabs v3 on the arena at a fraction of its list price.
Google Gemini TTS and Gemini 3.8 Live
Gemini 3.1 Flash TTS is still labeled preview. It costs $1 per 1M text input tokens and $20 per 1M audio output tokens (pricing) and sits at #8 on the arena. The older Gemini 2.5 Flash TTS ($0.50 / $10) and 2.5 Pro TTS ($1 / $20) are also still preview. Gemini 3.8 Live (September 15) is Google's realtime conversation model. It supports 97 languages and tops the speech-to-speech index. On Google's paid API tier, prompts and outputs are not used to improve Google's products (Gemini API terms).
OpenAI: GPT-Live-1, GPT-Realtime 2.1, gpt-4o-mini-tts
- GPT-Live-1: best for conversation. $0.05/min plus the reasoning model, 12 new preset voices, no custom voice cloning.
- gpt-4o-mini-tts ($0.60 per 1M text tokens in, $12 per 1M audio tokens out) and tts-1 ($15 per 1M characters): OpenAI's plain TTS options.
We found no newer OpenAI standalone TTS model as of September 23 (pricing).
MiniMax Speech 2.8, Fish Audio S2.1 Pro, Hume Octave 2, Deepgram Aura-2
- MiniMax Speech 2.8: HD at $100 per 1M characters, Turbo at $60. Cloning from a 10-second sample costs $1.50 per voice (pricing).
- Fish Audio S2.1 Pro: 83 languages, about 90 ms to first audio, $15 per 1M bytes. A
s2.1-pro-freeAPI tier is free until November 30, 2026, but businesses with more than $1M in annual revenue must contact Fish Audio before using it (pricing). - Hume Octave 2: built for emotional delivery. Cloning from a 15-second sample, 11 languages, $0.05–$0.15 per 1K characters. Its docs still call it preview (pricing).
- Deepgram Aura-2: $0.030 per 1K characters, and $0.075/min as a full voice agent (pricing).
Deepgram Aura-2, stock voice "Thalia". Generated through the API on July 13, 2026, from the same script as the ElevenLabs sample.
Microsoft MAI-Voice-2 and Amazon Nova 2 Sonic
MAI-Voice-2 (June 2, 2026) supports consent-gated zero-shot cloning from 5–60 seconds of audio. Microsoft's own pages disagree on its language count (9 to 17) and on whether it is generally available; the Foundry catalog still says "Preview". Amazon Nova 2 Sonic is AWS's speech-to-speech model. Choose either for procurement and governance fit inside Azure or AWS, not because it wins a listening test. We could not verify per-character prices on either vendor's pages.
Open weights: Qwen3-TTS, Kokoro, Chatterbox, Voxtral
| Model | License | Notes |
|---|---|---|
| Qwen3-TTS 0.6B / 1.7B | Apache 2.0 | 10 languages, 3-second voice cloning |
| Kokoro-82M | Apache 2.0 | Tiny, fast, 8 languages, no cloning |
| Chatterbox Multilingual v3 (Resemble) | MIT | 25 languages, cloning, watermark on all output |
| Voxtral-4B-TTS (Mistral) | CC BY-NC 4.0 | Non-commercial weights; commercial use via API at $0.016 / 1K chars |
PlayHT is gone. Meta acqui-hired the team in July 2025, and the platform shut down on December 31, 2025, according to migration notices from competitors; we found no primary Meta announcement. Sesame has no public API.
The voice is one piece of the job. In Teamday you write the script with an AI teammate, generate the voiceover in Audio, add music, and keep every version with the rest of the campaign, with no separate accounts to juggle.
Open Teamday Audio →How to Choose a Voice Model for Your Business
- Name the job. Recorded voiceover (TTS) and live conversation (speech-to-speech) are different products with different leaders. Pick from the right table above.
- Run your own script. Test 60–90 seconds with your product names, numbers, URLs, and prices, in every language you ship. Have a native speaker judge each language.
- Check the rights for your plan. Commercial use normally needs a paid plan. For a cloned voice, keep a written consent record from the speaker. For open weights, read the license: Voxtral's weights are non-commercial.
- Price per approved minute. Count retakes and fixes, not only the list price.
- Keep provenance. Store the provider, model ID, voice ID, script version, generation date, and approver with every published file. When a model is retired or a customer asks what they heard, you can answer and regenerate.
Which Voice Models Can You Use in Teamday?
| Where in Teamday | Models | Access |
|---|---|---|
| Audio studio (script → voiceover file) | Gemini 2.5 Flash TTS (default), Gemini 2.5 Pro TTS | Your connected Google Gemini account or Teamday AI balance |
| Voice conversations with agents | OpenAI GPT-Realtime 2.1; ElevenLabs Flash v2.5, v3, Multilingual v2 with Scribe v2 transcription; Deepgram Aura-2 with Nova-3 transcription | Your own OpenAI, ElevenLabs, or Deepgram key |
| Video narration in Reel's pipeline | ElevenLabs Flash v2.5 | Your own ElevenLabs key |
The Audio studio currently uses one preset voice. Add delivery directions such as "warm, unhurried" to the script. Teamday does not yet offer Cartesia, GPT-Live-1, Gemini 3.1 Flash TTS, Gemini 3.8 Live, or voice cloning. Use those directly with the vendor if you need them now.
Frequently Asked Questions
What is the best AI voice model in September 2026?
On the Artificial Analysis TTS arena (blind listener votes, checked September 23, 2026), Cartesia Sonic 3.6 ranks first, followed by Alibaba Qwen-Audio-3.0-TTS-Plus and Inworld Realtime TTS-2. For finished voiceovers, ElevenLabs v3 remains the most complete product (70+ languages, voice cloning, a large voice library). For live voice agents, OpenAI GPT-Live-1 and Gemini 3.8 Live lead the speech-to-speech rankings.
How much does AI voiceover cost per minute?
Roughly $0.01 to $0.10 per finished minute at API list prices, assuming about 900 characters per spoken minute. ElevenLabs v3 is $0.10 per 1,000 characters (about $0.09 per minute), ElevenLabs Flash and v3 Conversational $0.05 per 1,000, Inworld TTS-2 $25 per million characters, Deepgram Aura-2 $0.030 per 1,000, and Fish Audio S2.1 Pro $15 per million bytes. Retakes and edits usually cost more than the generation itself.
What is the difference between TTS and realtime speech-to-speech?
Text-to-speech reads a script you supply and is what you need for voiceovers, podcasts, and ads. Realtime speech-to-speech listens, reasons, and speaks in a live conversation, so interruption handling and turn latency matter as much as the voice. OpenAI GPT-Live-1 and Gemini 3.8 Live are speech-to-speech; ElevenLabs v3, Cartesia Sonic and Gemini TTS are TTS.
Can I use AI voices commercially?
Usually yes on paid plans, but check the provider's terms for your plan, the specific stock voice, and consent for any cloned voice. Open-weight models differ: Kokoro-82M and Qwen3-TTS are Apache 2.0, while Mistral's Voxtral TTS weights are CC BY-NC 4.0 (non-commercial; commercial use goes through Mistral's API).
Which voice models can I use in Teamday?
Teamday's Audio studio turns scripts into voiceovers with Google Gemini 2.5 Flash TTS or Gemini 2.5 Pro TTS, using your connected Gemini account or Teamday AI balance. Voice conversations with agents support OpenAI GPT-Realtime 2.1, ElevenLabs Flash v2.5, v3 and Multilingual v2, and Deepgram Aura-2 with your own provider keys. Video narration in Reel's pipeline uses ElevenLabs Flash v2.5.
Test it on your own script: one paragraph in, a voiceover file out, and you judge the result instead of a demo reel.
Generate a voiceover from your script →Related Guides
- Best AI music generators in September 2026
- Best AI video models
- Best frontier AI models
- Best AI avatar models
- Reel, Teamday's video producer
- Teamday model catalog
Research cutoff: September 23, 2026. Rankings are from Artificial Analysis. Prices and model status come from each vendor's pricing page or documentation on that date. Teamday availability was checked against the current Teamday code. Prices and preview status change often, so confirm them before committing a production budget.
