Best AI Voice Models in September 2026: Ranked, With Prices
Teamday· 16 min read· 2026-02-05· Updated 2026-09-23
AI VoiceTTSRealtime VoiceElevenLabsCartesiaOpenAIGemini2026

Best AI Voice Models in September 2026: Ranked, With Prices

Updated September 23, 2026. Three things changed since our July edition. Cartesia Sonic 3.6 (August 27) took first place on the independent Artificial Analysis TTS arena. OpenAI GPT-Live-1 shipped as an API at a flat $0.05 per minute. Google Gemini 3.8 Live (September 15) now leads the speech-to-speech rankings. ElevenLabs is no longer the quality leader on blind votes, but it is still the most complete voiceover product.

This guide is for founders and marketers who need voiceovers, podcast intros, ads, and explainer narration, plus teams building a voice agent. Picks come first, then the ranking, prices, and how to choose.

Need a voiceover today, not a vendor evaluation? Paste your script into Teamday Audio and get a downloadable voiceover from Gemini TTS, kept next to the video and files it belongs to.

Create your first voiceover →

Quick Answer: The Best AI Voice Model for Each Job

Your jobStart withWhy
Ad, explainer, or brand voiceoverElevenLabs v3Expressive delivery, 70+ languages, voice cloning, mature editor; $0.10 per 1K characters
Highest blind-test quality via APICartesia Sonic 3.6#1 on the Artificial Analysis TTS arena (ELO 1273); under 90 ms latency claim
Best quality per dollarInworld Realtime TTS-2 or Qwen-Audio-3.0-TTS#3 and #2 on the arena at a fraction of ElevenLabs v3's list price
Cheap high-volume narrationFish Audio S2.1 Pro or Deepgram Aura-2$15 per 1M bytes (free tier until Nov 30, 2026) / $0.030 per 1K characters
Live voice agent on the phone or webOpenAI GPT-Live-1 or Gemini 3.8 LiveTop of the speech-to-speech index; GPT-Live-1 is $0.05/min plus the reasoning model
Voiceovers inside Google CloudGemini 3.1 Flash TTS (preview)#8 on the arena, $1 text in / $20 audio out per 1M tokens
Self-hosted, commercial-safeQwen3-TTS or Kokoro-82MApache 2.0 weights; Qwen3-TTS clones from 3 seconds of audio

Do not choose from a 20-second demo. Test your own script, your product names, and every target language.

September 2026 Ranking: Text-to-Speech Quality

Artificial Analysis runs blind A/B listening tests and converts the votes into ELO scores. It is the best public quality signal, but it measures short English clips, not your brand voice, pronunciation of your product names, or long-form consistency.

RankModelVendorArena ELOReleasedVendor list price
1Sonic 3.6Cartesia1273Aug 2026Plans from $5/mo (100K credits)
2Qwen-Audio-3.0-TTS-PlusAlibaba1259Jul 2026API; AA lists $27.60 / 1M chars
3Realtime TTS-2Inworld1245Aug 2026$25 / 1M chars on demand
4Simba 3.2Speechify1237Jul 2026AA lists $6.60 / 1M chars
5Luna TTSVUI Labs1230Jun 2026AA lists $15 / 1M chars
6Realtime TTS-2 FlashInworld1210Aug 2026AA lists $10.40 / 1M chars
8Gemini 3.1 Flash TTSGoogle1199Apr 2026$1 text in / $20 audio out per 1M tokens
10v3 ConversationalElevenLabs1196Aug 2026$0.05 / 1K chars
11Sonic 3.5Cartesia1185May 2026Same plans as 3.6
14Speech 2.8 HDMiniMax1168Feb 2026$100 / 1M chars
15Eleven v3ElevenLabs1167Feb 2026$0.10 / 1K chars

Source: Artificial Analysis TTS leaderboard, checked September 23, 2026. Ranks 7, 9, 12, and 13 are BreezeBlue Breeze TTS 2, StepFun StepAudio 2.5, Smallest.ai Lightning V3.1 Pro, and Soniox TTS v2. Vendor prices are from each vendor's pricing page. Where a vendor does not publish a per-character rate, we show Artificial Analysis's normalized price and label it "AA lists".

Realtime Voice Agents: Speech-to-Speech Ranking

RankModelIndex scorePrice
1Gemini 3.8 Live Extended Thinking (High)82.6$0.005/min audio in, $0.018/min audio out
2OpenAI GPT-Live-181.5$0.05/min, billed per second, plus the backend reasoning model
3xAI Grok Voice Think Fast 2.081.3$0.08/min

Source: Artificial Analysis speech-to-speech, OpenAI GPT-Live-1 model page, Gemini API pricing, xAI pricing, all checked September 23, 2026.

GPT-Live-1 is a new pricing model. It is a full-duplex voice layer that listens and speaks at the same time and hands reasoning and tool calls to a backend model you pick. It runs only on the new v1/live/sessions endpoint. The older GPT-Realtime 2.1 (July 2026) is still offered at $32 per 1M audio input tokens and $64 per 1M audio output tokens. A cheaper mini version costs $10 and $20.

Price Per Finished Minute

Most TTS vendors price per character. A spoken minute at a normal pace of about 150 words is roughly 900 characters. The table below converts list prices on that assumption. It is an estimate, not a vendor quote.

ModelList price≈ Per spoken minute
ElevenLabs v3 / Multilingual v2$0.10 / 1K chars~$0.09
MiniMax Speech 2.8 HD$100 / 1M chars~$0.09
ElevenLabs Flash v2.5 / v3 Conversational$0.05 / 1K chars~$0.045
Deepgram Aura-2$0.030 / 1K chars~$0.027
Inworld Realtime TTS-2$25 / 1M chars~$0.023
Mistral Voxtral TTS (API)$0.016 / 1K chars~$0.014
Fish Audio S2.1 Pro$15 / 1M UTF-8 bytes~$0.014 (English)
OpenAI tts-1$15 / 1M chars~$0.014
Gemini TTS, gpt-4o-mini-ttsPriced per audio tokenMeasure with a real script

The generation fee is rarely what drives the cost. Retakes, pronunciation fixes, and review time usually cost more. Track total spend divided by approved finished minutes, not the list price.

The Models, One by One

Cartesia Sonic 3.6: best blind-test quality

Sonic 3.6 (August 27, 2026) is #1 on the TTS arena. Cartesia claims under 90 ms latency and covers 44 languages (61 locales). Instant voice cloning starts on the Pro plan and professional cloning on Startup. Pricing is credit-based: Pro is $5/month for 100K credits, Startup $49 for 1.25M, and Scale $299 for 8M (pricing). Built for streaming agents, it is also the strongest pure-quality pick for narration if your team is comfortable working through an API.

ElevenLabs v3: best complete voiceover product

ElevenLabs lists four TTS tiers:

  • v3: expressive, 70+ languages, $0.10 per 1K characters.
  • v3 Conversational: a low-latency v3 for agents, $0.05.
  • Multilingual v2: consistent long-form narration, $0.10.
  • Flash v2.5: about 75 ms, 32 languages, $0.05.

It no longer wins blind votes. It remains the easiest end-to-end product for a marketer: a voice library, cloning, an editor, dubbing, and music under one account.

ElevenLabs Flash v2.5, stock voice "Sarah". Generated through the API on July 13, 2026.

Alibaba Qwen-Audio-3.0-TTS and Inworld TTS-2: best quality per dollar

Qwen-Audio-3.0-TTS (July 21, 2026) comes in Plus and Flash versions. It covers 16 languages, supports voice cloning, and Alibaba claims about 300 ms first-packet latency for Flash. It is API-only, not open weights. Inworld's Realtime TTS-2 lists $25 per 1M characters on demand, dropping to $5 on Enterprise (pricing), and ranks #3. Both beat ElevenLabs v3 on the arena at a fraction of its list price.

Google Gemini TTS and Gemini 3.8 Live

Gemini 3.1 Flash TTS is still labeled preview. It costs $1 per 1M text input tokens and $20 per 1M audio output tokens (pricing) and sits at #8 on the arena. The older Gemini 2.5 Flash TTS ($0.50 / $10) and 2.5 Pro TTS ($1 / $20) are also still preview. Gemini 3.8 Live (September 15) is Google's realtime conversation model. It supports 97 languages and tops the speech-to-speech index. On Google's paid API tier, prompts and outputs are not used to improve Google's products (Gemini API terms).

OpenAI: GPT-Live-1, GPT-Realtime 2.1, gpt-4o-mini-tts

  • GPT-Live-1: best for conversation. $0.05/min plus the reasoning model, 12 new preset voices, no custom voice cloning.
  • gpt-4o-mini-tts ($0.60 per 1M text tokens in, $12 per 1M audio tokens out) and tts-1 ($15 per 1M characters): OpenAI's plain TTS options.

We found no newer OpenAI standalone TTS model as of September 23 (pricing).

MiniMax Speech 2.8, Fish Audio S2.1 Pro, Hume Octave 2, Deepgram Aura-2

  • MiniMax Speech 2.8: HD at $100 per 1M characters, Turbo at $60. Cloning from a 10-second sample costs $1.50 per voice (pricing).
  • Fish Audio S2.1 Pro: 83 languages, about 90 ms to first audio, $15 per 1M bytes. A s2.1-pro-free API tier is free until November 30, 2026, but businesses with more than $1M in annual revenue must contact Fish Audio before using it (pricing).
  • Hume Octave 2: built for emotional delivery. Cloning from a 15-second sample, 11 languages, $0.05–$0.15 per 1K characters. Its docs still call it preview (pricing).
  • Deepgram Aura-2: $0.030 per 1K characters, and $0.075/min as a full voice agent (pricing).

Deepgram Aura-2, stock voice "Thalia". Generated through the API on July 13, 2026, from the same script as the ElevenLabs sample.

Microsoft MAI-Voice-2 and Amazon Nova 2 Sonic

MAI-Voice-2 (June 2, 2026) supports consent-gated zero-shot cloning from 5–60 seconds of audio. Microsoft's own pages disagree on its language count (9 to 17) and on whether it is generally available; the Foundry catalog still says "Preview". Amazon Nova 2 Sonic is AWS's speech-to-speech model. Choose either for procurement and governance fit inside Azure or AWS, not because it wins a listening test. We could not verify per-character prices on either vendor's pages.

Open weights: Qwen3-TTS, Kokoro, Chatterbox, Voxtral

ModelLicenseNotes
Qwen3-TTS 0.6B / 1.7BApache 2.010 languages, 3-second voice cloning
Kokoro-82MApache 2.0Tiny, fast, 8 languages, no cloning
Chatterbox Multilingual v3 (Resemble)MIT25 languages, cloning, watermark on all output
Voxtral-4B-TTS (Mistral)CC BY-NC 4.0Non-commercial weights; commercial use via API at $0.016 / 1K chars

PlayHT is gone. Meta acqui-hired the team in July 2025, and the platform shut down on December 31, 2025, according to migration notices from competitors; we found no primary Meta announcement. Sesame has no public API.

The voice is one piece of the job. In Teamday you write the script with an AI teammate, generate the voiceover in Audio, add music, and keep every version with the rest of the campaign, with no separate accounts to juggle.

Open Teamday Audio →

How to Choose a Voice Model for Your Business

  1. Name the job. Recorded voiceover (TTS) and live conversation (speech-to-speech) are different products with different leaders. Pick from the right table above.
  2. Run your own script. Test 60–90 seconds with your product names, numbers, URLs, and prices, in every language you ship. Have a native speaker judge each language.
  3. Check the rights for your plan. Commercial use normally needs a paid plan. For a cloned voice, keep a written consent record from the speaker. For open weights, read the license: Voxtral's weights are non-commercial.
  4. Price per approved minute. Count retakes and fixes, not only the list price.
  5. Keep provenance. Store the provider, model ID, voice ID, script version, generation date, and approver with every published file. When a model is retired or a customer asks what they heard, you can answer and regenerate.

Which Voice Models Can You Use in Teamday?

Where in TeamdayModelsAccess
Audio studio (script → voiceover file)Gemini 2.5 Flash TTS (default), Gemini 2.5 Pro TTSYour connected Google Gemini account or Teamday AI balance
Voice conversations with agentsOpenAI GPT-Realtime 2.1; ElevenLabs Flash v2.5, v3, Multilingual v2 with Scribe v2 transcription; Deepgram Aura-2 with Nova-3 transcriptionYour own OpenAI, ElevenLabs, or Deepgram key
Video narration in Reel's pipelineElevenLabs Flash v2.5Your own ElevenLabs key

The Audio studio currently uses one preset voice. Add delivery directions such as "warm, unhurried" to the script. Teamday does not yet offer Cartesia, GPT-Live-1, Gemini 3.1 Flash TTS, Gemini 3.8 Live, or voice cloning. Use those directly with the vendor if you need them now.

Frequently Asked Questions

What is the best AI voice model in September 2026?

On the Artificial Analysis TTS arena (blind listener votes, checked September 23, 2026), Cartesia Sonic 3.6 ranks first, followed by Alibaba Qwen-Audio-3.0-TTS-Plus and Inworld Realtime TTS-2. For finished voiceovers, ElevenLabs v3 remains the most complete product (70+ languages, voice cloning, a large voice library). For live voice agents, OpenAI GPT-Live-1 and Gemini 3.8 Live lead the speech-to-speech rankings.

How much does AI voiceover cost per minute?

Roughly $0.01 to $0.10 per finished minute at API list prices, assuming about 900 characters per spoken minute. ElevenLabs v3 is $0.10 per 1,000 characters (about $0.09 per minute), ElevenLabs Flash and v3 Conversational $0.05 per 1,000, Inworld TTS-2 $25 per million characters, Deepgram Aura-2 $0.030 per 1,000, and Fish Audio S2.1 Pro $15 per million bytes. Retakes and edits usually cost more than the generation itself.

What is the difference between TTS and realtime speech-to-speech?

Text-to-speech reads a script you supply and is what you need for voiceovers, podcasts, and ads. Realtime speech-to-speech listens, reasons, and speaks in a live conversation, so interruption handling and turn latency matter as much as the voice. OpenAI GPT-Live-1 and Gemini 3.8 Live are speech-to-speech; ElevenLabs v3, Cartesia Sonic and Gemini TTS are TTS.

Can I use AI voices commercially?

Usually yes on paid plans, but check the provider's terms for your plan, the specific stock voice, and consent for any cloned voice. Open-weight models differ: Kokoro-82M and Qwen3-TTS are Apache 2.0, while Mistral's Voxtral TTS weights are CC BY-NC 4.0 (non-commercial; commercial use goes through Mistral's API).

Which voice models can I use in Teamday?

Teamday's Audio studio turns scripts into voiceovers with Google Gemini 2.5 Flash TTS or Gemini 2.5 Pro TTS, using your connected Gemini account or Teamday AI balance. Voice conversations with agents support OpenAI GPT-Realtime 2.1, ElevenLabs Flash v2.5, v3 and Multilingual v2, and Deepgram Aura-2 with your own provider keys. Video narration in Reel's pipeline uses ElevenLabs Flash v2.5.

Test it on your own script: one paragraph in, a voiceover file out, and you judge the result instead of a demo reel.

Generate a voiceover from your script →

Research cutoff: September 23, 2026. Rankings are from Artificial Analysis. Prices and model status come from each vendor's pricing page or documentation on that date. Teamday availability was checked against the current Teamday code. Prices and preview status change often, so confirm them before committing a production budget.