Google's Flash TTS models let creators design voices from text descriptions — 100+ languages, stage directions, two-voice dialogue from one script.
Google Gemini Flash TTS: Design AI Voices from Text Descriptions
Google has launched Gemini 3.8 Flash TTS and Flash-Lite TTS, two new text-to-speech models that let creators design custom voices from scratch using plain text descriptions. Available now through the Gemini API and Google AI Studio, these models support more than 100 languages and introduce features like stage directions per line, two-voice dialogue from a single script, and voice cloning from a 30-second audio sample.
Design Any Voice You Can Describe
The standout feature of Flash TTS is generating new voices by describing them in natural language. Instead of picking from pre-built voice banks, you can type something like "a warm, authoritative male voice with a slight British accent, speaking slowly and clearly" and the model generates it instantly. This opens voice design to anyone who can describe what they hear in their head — no studio session or voice actor booking required.
For Hong Kong creators producing bilingual and multilingual content, this means you can design Cantonese, Mandarin, and English voices that match your brand's personality across markets. The 100+ language support covers Hong Kong's key audiences, and voice characteristics carry over naturally when switching between languages.
Stage Directions and Two-Voice Dialogue
Flash TTS introduces stage directions — inline instructions that control delivery for individual lines in your script. You can add [whisper], [excited], [slowly], or [commanding] directly into the text and the model adjusts tone, pace, and emotion per sentence. Combined with two-voice dialogue generation from a single script, this is a powerful tool for podcasts, audiobooks, explainer videos, and character-driven social content.
Both models also handle nonverbal sounds like laughter and sighs, adding natural texture to AI-generated voiceovers that previously required manual audio editing.
Two Tiers for Different Workloads
Gemini 3.8 Flash TTS is the premium model built for quality-critical use cases like podcasts, audiobooks, and game characters, with full voice description and voice cloning capabilities. Flash-Lite TTS is optimised for cost-efficient speech generation at scale — ideal for dubbing, social media content, and voice agent applications.
Pricing is billed per million tokens, with text input at $0.50-1.00 and audio output varying by model tier. A free tier is available through Google AI Studio, giving HK creators a risk-free way to test voice designs before committing to production.
What This Means for HK Creators
For Hong Kong's creator economy, Flash TTS solves two recurring problems: voice consistency across languages and rapid voice prototyping for ad campaigns. If you're producing a Cantonese explainer video that needs a specific brand voice, then repurposing it for an English-speaking audience, Flash TTS can maintain character while switching languages naturally.
The model also pairs naturally with Cooly Studio for end-to-end AI video production — generate your script, design the voice, sync it to video, all in one workflow without leaving your production pipeline.
Getting Started
Flash TTS models are rolling out now through the Gemini API and Google AI Studio, with Gemini Enterprise API access following soon. The free tier lets you experiment with voice design and stage directions, so there's no upfront cost to evaluate whether Flash TTS fits your production needs.
Frequently Asked Questions
Q: Can I use Flash TTS for Cantonese voice generation? A: Yes. Flash TTS supports 100+ languages including Cantonese and Mandarin, with natural cross-language voice cloning.
Q: How does voice cloning work with a 30-second sample? A: Upload a 30-second audio clip through the Gemini API and Flash TTS builds a matching voice profile in seconds — no lengthy training required.
Q: What's the difference between Flash TTS and Flash-Lite TTS? A: Flash TTS is the premium model with full voice description and cloning for quality-critical content. Flash-Lite TTS is optimised for low-cost, high-volume speech generation.
Q: Can I use Flash TTS voices in Cooly Studio? A: Yes. Flash TTS can be integrated into your Cooly Studio video production pipeline for AI voiceovers, narration, and bilingual content projects.
Q: Is there a free tier to test Flash TTS? A: Yes. Google offers free tier access through Google AI Studio. Paid pricing starts at $0.50 per million input tokens.
