Automate Basics
New model · Google

Google brings new Gemini speech models to Notebook and Vids

By Automate Basics, written with AI from Google's original

· 2 min read

Hand holding a marked-up script beside a microphone and headphones
Image: AI-generated illustration.

Google has introduced two Gemini text-to-speech models for creating and directing spoken audio. Gemini 3.8 Flash TTS is rolling out in Gemini Notebook, while Flash-Lite TTS is rolling out in Google Vids.

What changed in Gemini text-to-speech

Google says its new Gemini 3.8 Flash TTS and Flash-Lite TTS models can turn scripts into speech with control over pacing, emotion and conversational delivery. Both models support line-by-line direction and dialogue between two distinct speakers. Google positions Flash TTS for designing character voices and directing performances, while Flash-Lite TTS is intended for higher-volume work such as dubbing and audio production.

Flash TTS can create voices from written descriptions or reproduce a voice from an audio sample, according to Google. Voice replication requires a matching verbal consent recording from the voice owner. Google also says audio generated by its Gemini Audio models carries a SynthID watermark. For a team making spoken content, the practical change is more control over how a script sounds, not just which voice reads it.

Who can use the new speech models?

Google says Gemini 3.8 Flash TTS is rolling out in Gemini Notebook for general users, and Flash-Lite TTS is rolling out in Google Vids. Developers can access both models through Google AI Studio and the Gemini API as the rollout begins. Access through the Gemini Enterprise API is coming later.

The announcement does not specify which paid plans include the models or what they cost. It also does not give a date for the planned Gemini Enterprise access. If you use Notebook or Vids at work, check whether the relevant model is available to you before planning a project around it. The two apps are getting different models, so the voice-design options Google describes for Flash TTS should not be assumed to appear in Vids.

How do I try Gemini speech at work?

A short training narration or internal explainer is a sensible first task for Gemini speech generation. Prepare a script you already have permission to use, then try the model available in your app. If the piece includes a conversation, Google says the models can handle two speakers and directions for how individual lines should be delivered.

Listen to the result before sharing it. Check names and specialist terms, whether the pace suits the audience, and whether each speaker stays distinct. For longer material, check that the voice remains consistent throughout. If you are considering voice replication, use only your own voice or one you have rights to use, and expect Google’s consent verification requirement. Treat the generated recording as a draft that needs editorial approval.

Frequently asked questions

What is the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS?

Google describes Flash TTS as the model for creating custom voices and giving detailed performance direction. Flash-Lite TTS is aimed at higher-volume speech work, including dubbing and voice agents. Both models support control over how lines are spoken.

Can I use Gemini's new speech models without the API?

Google says Flash TTS is rolling out in Gemini Notebook and Flash-Lite TTS is rolling out in Google Vids for general users. The announcement does not say which paid plans include them or give a price. Availability may depend on the rollout.

Tools in this piece

Written with AI from Google's original and published after automatic checks: every figure here appears in the original, and no sentence is copied from it. The picture is AI-generated. The original is the authority.

Source: Gemini 3.8 text-to-speech says hello, Google, 23 Sept 2026.