Gemini 3.8 text-to-speech says hello
3 hours ago
- Google introduced two new text-to-speech models: Gemini 3.8 Flash TTS for deep creative control and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient scale.
- Flash TTS allows creating bespoke voices from scratch using natural language prompts, supporting over 100 languages and dialects, with features like voice replication from a 30-second sample.
- Both models offer line-by-line performance direction, long-form generation with minimal drift, native two-speaker scene staging, and scripted vocal bursts.
- Gemini 3.8 Flash TTS achieved #1 on Hume AI’s Voice Design Benchmark and top spots on Voice Arena for multiple languages.
- Voice replication includes consent verification, SynthID watermarking, and C2PA credentials for trust and transparency.
- The models are available via Gemini API and Google AI Studio, with enterprise and consumer versions rolling out soon (Notebook LM, Google Vids).