- Start
- Sep 23, 202690% CONFIDENCEfrom the source
Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS Released
- Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026 as "our most expressive audio generation models yet": Flash TTS for creative direction and character design, Flash-Lite TTS for high-volume, cost-efficient dubbing and voice agents[1]
- Flash TTS creates new voices from natural-language prompts in more than 100 languages and dialects, offers more than 2,000 ready-made voices, and can replicate a voice from a 30-second sample once the owner's recorded verbal consent matches the reference speaker[1]
- Both take line-by-line direction on pacing, emotion and accent, keep a voice stable over hours of audio, stage two-speaker scenes from one script, and render cues such as <laughs>, <sigh> and backchannels like "mhm"[1]
- They rolled out in the Gemini API and Google AI Studio, with Flash TTS in Gemini Notebook and Flash-Lite TTS in Google Vids; Gemini Enterprise access is coming soon[1][5]
- The Gemini API changelog lists the release under 22 September, with a new Voices endpoint; Flash-Lite TTS replaces the Gemini 3.1 Flash TTS preview[2]
- AI Studio voice replication is blocked in the UK, the EEA, India, Texas and Illinois[5]
Notable features
- First overall on Hume AI's Voice Design Benchmark (71.4) and first in accent modelling (60.8); Flash and Flash-Lite take first and second on Hume's Overall Quality Index, though ElevenLabs scored higher in Hume's individual voice-quality categories[1][5]
- Flash TTS costs $0.50 per million text input tokens and $9.00 per million audio output tokens (about $0.00225 per 10 seconds) through 31 December 2026, doubling to $1.00 and $18.00 from 1 January 2027; Flash-Lite TTS costs $0.50 and $6.00, rising to $1.00 and $12.00[3]
- Text input of up to 8,192 tokens and up to 16,384 output tokens on the Gemini API[4]
- Every clip carries a SynthID watermark, and replicated voices also carry C2PA credentials[1][5]
- Partners integrating the models include Figma, HeyGen, Wondercraft, 99.co and Ollang[1]
References 687% CONFIDENCEOverall confidence: 87%How well the pin's source and references back up its dates.Weighted average of how firmly 6 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speechblog.google· Posted Sep 28, 2026· Starts Sep 23, 2026 ✓· 17% of score
Google's[2][3][4] post is dated "Sep 23, 2026" and says both models are "rolling out starting today"; The Next Web[6] reports "Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Wednesday". The Gemini API changelog files the GA a day earlier, under September 22.
- [2]85%Release notes | Gemini APIai.google.dev· Added Sep 28, 2026· 17% of score
The Gemini API changelog under September 22, 2026 lists both models as generally available with the new Voices endpoint, voice design, voice replication and an extended library of 150+ voices, and says Flash-Lite TTS replaces gemini-3.1-flash-tts-preview.[1]
- [3]85%Gemini Developer API pricingai.google.dev· Added Sep 28, 2026· 17% of score
Flash TTS costs $0.50 per million text input tokens and $9.00 per million audio output tokens through 31 December 2026, then $1.00 and $18.00; Flash-Lite TTS $0.50 and $6.00, then $1.00 and $12.00.[1]
- [4]85%Gemini 3.8 Flash TTS | Gemini APIai.google.dev· Added Sep 28, 2026· 17% of score
The API model page for gemini-3.8-flash-tts: text in, audio out, an 8,192-token input limit and 16,384 output tokens (Gemini API serving limit), latest update September 2026.[1]
- [5]85%Gemini can now clone your voice and perform scripts like an actorandroidauthority.com· Published Sep 24, 2026· 16% of score
Android Authority on the rollout to Gemini Notebook and Google[2][3][4] Vids; it notes ElevenLabs still scored higher in Hume's individual voice-quality and human-like-variation categories, and that AI Studio voice replication is blocked in the UK, the EEA, India, Texas and Illinois.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.