Updates·September 24, 2026, 05:12
Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS
AI-generated and checked against the sources listed below.
Google has released two new speech AI models that can read text aloud with natural voices, design entirely new voices from a text description and clone a voice from just 30 seconds of audio.

What's new
On September 23, 2026, Google launched two new speech AIs: Gemini 3.8 Flash TTS and its cheaper little sibling Gemini 3.8 Flash-Lite TTS. TTS stands for "text-to-speech," meaning AI that turns written text into spoken audio. The new models read text aloud with a natural voice and can be controlled down to the smallest detail, such as pauses, tone of voice and sighs.
What's clever
The cleverest part is that you can design an entirely new voice from a text description. Write, for example, "a warm, slightly husky female voice with a Scottish accent," and the AI comes up with the voice. You can also clone a real voice from just 30 seconds of audio recording, though only with verbal consent from the person. Google offers more than 2,000 ready-made voices and supports more than 100 languages, including regional variants such as Mexican Spanish and Quebec French. The models can also handle two voices in dialogue and add non-verbal sounds such as laughter and sighs, without the voice "drifting" or changing character over the course of long recordings. All clips are invisibly marked with Google's SynthID watermark, so you can tell the audio is AI-generated.
Cheaper or better?
Flash TTS costs $0.50 per million text tokens (tokens are the chunks of text the AI computes in) and $9 per million audio tokens through the end of 2026, after which the price doubles to $1 and $18 respectively from January 1, 2027. Flash-Lite TTS has the same price for text but only $6 per million audio tokens, making it the cheap model for large volumes of speech, such as dubbing. In practice, this corresponds to about $0.81 per hour of audio with Flash and $0.54 with Flash-Lite. According to Google, the models top both Hume AI's quality measurement and a Voice Design Benchmark with a score of 71.4, ahead of competitors, according to Google's own figures. An independent comparison with, for example, OpenAI or ElevenLabs does not yet exist.
What it's good at
Flash TTS is made for creative content: podcasts, audiobooks and video game characters, where you want acting and emotion in the voice. Flash-Lite TTS is built for companies that need to generate large amounts of speech cheaply, for example for voice agents, video dubbing or automatic reading of content. The models are rolling out via the Gemini API and Google AI Studio now and will soon come to Gemini Enterprise as well as Gemini Notebook and Google Vids.
Sources
Get the week's AI news in your inbox
Choose your level, topics and length. One email a week, unsubscribe at any time.
Subscribe to Promptly Newsletter



