Skip to content
AI.info

The Pulse

Google Launches Gemini 3.8 Flash TTS for Custom Voices

Google is releasing Gemini 3.8 Flash TTS and Flash-Lite TTS, offering custom voice design, line-by-line performance controls, dialogue staging and built-in watermarking.

Google Launches Gemini 3.8 Flash TTS for Custom Voices

AI.info Team ·

Google puts voice design inside Gemini

Google is releasing two text-to-speech models that let users design voices, direct performances line by line and generate dialogue scenes from a single script. Gemini 3.8 Flash TTS targets detailed creative control, while Gemini 3.8 Flash-Lite TTS is aimed at higher-volume uses such as dubbing, audio production and voice agents.

The models begin rolling out on September 23, 2026, through the Gemini API and Google AI Studio. Gemini 3.8 Flash TTS is also coming to Gemini Enterprise and is available in Gemini Notebook, while the Flash-Lite version is rolling out to Google Vids for general users.

Google describes the pair as its most expressive audio generation models yet. The company says they support more than 100 languages and dialects, with controls for accents, pacing, emotion, vocal mannerisms and conversational timing.

“Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio.”
— Leland Rechis, Group Product Manager, and Alan Cowen, Director, Research Science, on behalf of the Gemini Audio Team, Google

From preset voices to generated characters

Gemini 3.8 Flash TTS can create a voice from a natural-language description rather than requiring developers to select from a fixed catalog. Users can specify a character’s role, accent and vocal qualities, then save the resulting profile for use across a project.

Google says the model provides access to more than 2,000 production-ready voices and supports regional varieties including Mexican Spanish, Quebec French and Scots English. The company also positions the system for original character work in games, immersive audiobooks, podcasts and interactive media.

Voice replication is supported from a 30-second audio sample, provided the user has permission to use the source voice. Google says the process includes consent verification, SynthID watermarking and C2PA credentials. Voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland or India.

Line-by-line direction and two-speaker scenes

Both models accept performance instructions alongside the script. A creator can specify a calm delivery, a whispered line, a change in emotional intensity or a regional pronunciation for individual passages instead of applying one style to an entire recording.

The models also support nonverbal cues and conversational backchanneling. Google lists examples such as laughter, sighs, gasps and short active-listening responses including “mhm” and “yeah.” Gemini 3.8 Flash TTS can stage two-speaker conversations from one script while keeping the voices distinct and managing turn-taking.

Long-form generation is another target. Google says the models maintain voice quality, pacing and character timbre across hours of audio with limited speaker drift, a capability aimed at audiobooks, podcasts and other extended recordings.

Google cites Hume and Voice Arena results

Google says Gemini 3.8 Flash TTS ranks first on Hume AI’s Voice Design Benchmark with a score of 71.4 and leads its accent-modeling evaluation with a score of 60.8. The company also says Flash TTS and Flash-Lite TTS take the first and second positions, respectively, on Hume AI’s Overall Quality Index.

Google reports that the models occupy top positions in blind human-preference evaluations from Voice Arena across Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. Those results are company-reported benchmark claims rather than an independent assessment included in the announcement.

Watermarking is built into every generated clip

Google says every audio clip produced by its Gemini Audio models carries an imperceptible SynthID watermark. The watermark is designed to help identify AI-generated speech, while C2PA credentials are used with voice replication to provide additional information about the media’s origin.

Developers can begin testing the models in Google AI Studio and through the Gemini API. Google lists Agora, LiveKit, Pipecat and Vercel among platforms using the API, and names Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as companies integrating the new TTS systems.

Source

Google

Explore

More articles