Google Launches Gemini 3.8 Flash and Flash-Lite TTS
Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, targeting different speech-generation workloads. Flash is designed for character creation and detailed scene direction, while Flash-Lite is optimized for cost-efficient, high-volume dubbing, audio production and expressive voice agents.
Gemini 3.8 Flash TTS can generate a voice from natural-language descriptions of its role, accent and other characteristics. Google says it covers more than 100 languages and dialects and provides access to over 2,000 production-ready voices, including Mexican Spanish, Quebec French and Scots English options.
Both models let users direct individual lines with script cues. Google specifies granular control over acting cues, pacing, dialect shifts and backchanneling for Flash TTS; Flash-Lite offers fine-grained control over tone, pacing and expressive nuance. A single script can also stage a two-speaker conversation while keeping the voices separate and managing natural turn-taking.
For longer productions, Google claims the models can maintain voice quality, natural pacing and character timbre across hours of continuous audio with minimal speaker drift. Scripts may include nonverbal reactions such as laughter, sighs and gasps, as well as active-listening interjections such as “mhm” or “yeah.”
Flash TTS can replicate a vocal profile from a 30-second sample of the user’s voice or another voice they have rights to use. Before a profile is created, the system requires a verbal consent recording from the voice owner that matches the reference speaker. Google says every clip generated by Gemini Audio models carries an imperceptible SynthID watermark; it also lists C2PA credentials among the protections for voice replication.
Rollouts for both models began in the Gemini API and Google AI Studio. Flash TTS is also rolling out in Gemini Notebook, while Flash-Lite is coming to Google Vids. Enterprise access through the Gemini Enterprise API is listed as coming soon. Google separately previewed voice remixing, which will allow adjustments to timbre, pitch, pace and accent, but did not provide a release date.
Practical context: The split gives developers a choice between deeper performance direction and efficient production at volume. The distinctions matter: Google explicitly assigns dialect-shift controls and voice replication to Flash TTS, not Flash-Lite, while replication also carries consent requirements and regional restrictions in AI Studio.
Google reports that Flash TTS scored 71.4 on Hume AI’s Voice Design Benchmark and 60.8 for accent modeling. It also says Flash and Flash-Lite ranked first and second, respectively, on Hume AI’s Overall Quality Index. These are vendor-reported results, and Google’s announcement does not provide enough methodological detail to assess them independently.
| Model | Primary purpose | Announced rollout |
|---|---|---|
| Gemini 3.8 Flash TTS | Voice and character creation with detailed scene direction | Gemini API, Google AI Studio and Gemini Notebook; Gemini Enterprise API coming soon |
| Gemini 3.8 Flash-Lite TTS | High-volume dubbing, audio production and expressive voice agents | Gemini API, Google AI Studio and Google Vids; Gemini Enterprise API coming soon |
Geographic restrictions on voice replication
Voice replication through Google AI Studio is unavailable in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India. Creating a profile requires a verbal consent recording from the voice owner that matches the reference speaker.
Sources
Event date: 2026-09-23. Primary source date: 2026-09-23.