Google has expanded its Gemini audio portfolio with new speech-generation models for developers and businesses. On September 23, 2026, the company introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Both models are available through the Gemini API and Google AI Studio.
The launch gives developers greater control over voice generation, character design, and conversational audio. Meanwhile, Google is targeting applications ranging from audiobooks and podcasts to dubbing and real-time voice agents.
The latest models build on Google’s September 15 release of Gemini 3.8 Live and Gemini 3.5 Transcribe. Consequently, developers can access speech generation, transcription, and real-time voice interaction within the broader Gemini audio ecosystem.
Gemini 3.8 Flash TTS Targets Expressive Voice Generation
Google designed Gemini 3.8 Flash TTS for creative audio production and detailed voice direction. The model allows developers to create custom voices through natural-language prompts.
For example, users can specify a character’s vocal identity, accent, and emotional tone. Additionally, they can direct dialogue line by line, adjusting pacing and delivery. This gives creators greater control over performances in interactive media.
The model supports applications such as gaming, immersive audiobooks, podcasts, and conversational experiences. Moreover, developers can generate dialogue for multiple speakers and shape the delivery of individual lines.
Google also supports voice design and voice replication through its speech-generation tools. However, voice replication requires consent verification. Users must provide a verbal consent recording from the voice owner that matches the reference speaker.
Flash-Lite TTS Focuses on Scalable Audio Production
Alongside Flash TTS, Google introduced Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. The model targets businesses that need to produce large amounts of audio.
In particular, Google highlights dubbing, audio content creation, and expressive voice agents as key applications. Furthermore, developers can control tone, pacing, and vocal expression while producing audio at scale.
Both models share the same API structure and prompting format. Therefore, developers can switch between them by changing the model parameter, depending on their performance and production requirements.
Google is also adding safeguards to its audio-generation technology. Every generated audio clip carries SynthID, an imperceptible watermark designed to help identify AI-generated speech. Consequently, the company aims to support transparency and reduce the risk of misuse.
Availability Across Google’s Developer Platforms
The new models began rolling out on September 23 through the Gemini API and Google AI Studio. Developers can test their capabilities in AI Studio before integrating them into applications.
Additionally, Google is extending access to other products. Gemini 3.8 Flash TTS is coming to Gemini Enterprise through its API, while Google Vids will support Flash-Lite TTS.
Google is also working with platforms including Agora, LiveKit, Pipecat, and Vercel. These integrations aim to help developers deploy speech-generation systems and voice interfaces. Meanwhile, partners such as Figma, HeyGen, Wondercraft, and Ollang are exploring applications in media localisation and conversational AI.
The launch expands Google’s efforts to make advanced speech technology accessible to developers and enterprises. Ultimately, the two models address different needs, from creative voice production to large-scale audio generation. Their adoption will depend on the quality, cost, and reliability developers achieve in real-world applications.








