Artificial intelligence

Google Launches Gemini 3.8 Flash TTS and Flash-Lite for Customizable Text-to-Speech

Google announced two new text-to-speech models that enable the creation of custom voices, line-by-line performance direction, and multi-speaker dialogues in more than 100 languages. Gemini 3.8 Flash TTS is available to developers through the Gemini API and Google AI Studio, while Flash-Lite targets large-scale audio applications.

2026-09-23
4 min read
73 views
certi.news Editorial Team
Google Launches Gemini 3.8 Flash TTS and Flash-Lite for Customizable Text-to-Speech

Google announced the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS text-to-speech models, focusing on creating more expressive voices and controlling how each line of text is delivered. The company says the two models target developers, creators, and organizations building audiobooks, podcasts, voice agents, and interactive experiences.

Two Models for Different Uses

Gemini 3.8 Flash TTS is designed to create custom voices and characters through natural-language prompts, with control over role, accent, and voice characteristics across more than 100 languages and dialects. Google also offers a library of more than 2,000 production-ready voices, including regional variations such as Mexican Spanish, Quebec French, and Scottish English.

Gemini 3.8 Flash-Lite TTS focuses on high-volume, lower-cost applications, such as dubbing, audio content production, and voice agents that require large-scale control over tone, rhythm, and expressive details.

Detailed Control Over Vocal Performance

The two models enable line-by-line direction of delivery using instructions that specify speed, emotion, accent, and performance style. Their capabilities include creating two-speaker dialogue scenes from a single script while maintaining distinct voices and organizing turn-taking, in addition to nonverbal audio effects such as laughter, sighs, gasps, and brief listening cues.

Google says Flash TTS can maintain voice quality and rhythm and character identity throughout long audio productions lasting hours, while reducing changes in the speaker’s characteristics. The company is also working on a voice remixing feature that will allow users to adjust pitch, speed, and accent, but it is still “coming soon.”

Voice Cloning and Announced Safeguards

A consistent voice profile can be created from a 30-second voice sample, provided that the sample belongs to the user or that the user has the rights to use it. Google requires a verbal consent recording from the voice owner and matches it with the reference speaker before creating the voice clone.

The company confirms that every audio clip produced by Gemini Audio models carries an invisible watermark through SynthID, alongside the use of C2PA credentials in voice-cloning capabilities. These mechanisms are intended to facilitate the detection of AI-generated audio and enhance transparency about its origin. However, consent to use another person’s voice and regulatory restrictions remain decisive factors outside the technical verification mechanism itself.

Availability and What Changes in Practice?

The rollout of Gemini 3.8 Flash TTS to developers has begun through the Gemini API and Google AI Studio, and to users through Gemini Notebook, with availability later through the API in Gemini Enterprise. Flash-Lite TTS has also begun rolling out through Google Vids. Google additionally announced partnerships or integrations with platforms and companies including Agora, LiveKit, Pipecat, Vercel, Figma, HeyGen, Wondercraft, and Ollang.

In practice, the announcement moves text-to-speech away from reliance on fixed voices toward a workflow closer to directing a vocal performance: designing a character, writing delivery instructions, and then producing long or multilingual dialogue. Google says Flash TTS ranked first in Hume AI’s Voice Design Benchmark with a score of 71.4 and in accent modeling with a score of 60.8, while the two models took first and second place in its overall quality index. These are results presented by the company as part of the announcement and do not replace independent evaluation of cost and performance in actual production scenarios.

Voice-cloning functionality is not available through AI Studio in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India, according to the restrictions mentioned in the announcement.

News source
Google Official News
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news