Logo
FrontierNews.ai

Google's New Text-to-Speech AI Can Clone Your Voice and Generate 2,000+ Unique Voices

Google DeepMind has released two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, that can generate speech in over 2,000 different voices across more than 100 languages and dialects, including Japanese. The advanced model ranks first on Google's Pronunciation Robustness Benchmark and second on the Provider Voice Arena Leaderboard, marking a significant step forward in how creators can produce audio content for games, audiobooks, podcasts, and other media.

What Makes These New Speech Models Different From Earlier Versions?

The two models serve different purposes in the audio creation workflow. Gemini 3.8 Flash TTS is designed for advanced, creative processing, allowing users to generate entirely new voices from scratch using natural language prompts. If you want a lower voice or a specific accent, you can simply describe what you're looking for, and the AI suggests appropriate options. Gemini 3.8 Flash-Lite TTS, by contrast, is optimized for high-volume processing, making it ideal for businesses and creators who need to generate large amounts of audio quickly.

Both models include a voice cloning feature that sets them apart from earlier text-to-speech tools. Users can record a 30-second audio sample and have text read aloud in their own voice, provided they give explicit consent. This consent requirement is built in as a security measure to prevent misuse, ensuring that audio from movies, anime, or other sources cannot be reproduced without permission.

How to Create Professional Audio Content With Gemini TTS?

  • Voice Generation: Use natural language prompts to describe the voice characteristics you want, such as pitch, tone, or accent, and let the AI suggest matching voices from its library of over 2,000 options.
  • Voice Cloning: Record a 30-second audio sample of your own voice or a speaker's voice (with their consent) and use it to generate text-to-speech output that maintains the speaker's unique characteristics.
  • Long-Form Audio Production: Generate continuous audio lasting several hours while maintaining high quality, natural pacing, and consistent character voice traits, including realistic expressions like laughter, sighs, and interjections.
  • Multi-Speaker Conversations: Create dialogue between multiple characters with distinct voices and emotional expressions, useful for audiobooks, podcasts, and interactive media.
  • Watermark Protection: All generated audio is automatically watermarked with SynthID, Google's audio authentication technology, to ensure transparency and prevent misuse.

How Does Performance Compare to Competing Models?

Google's new models achieved top-tier performance on multiple industry benchmarks. In the Hume AI voice benchmark, Gemini 3.8 Flash TTS scored the highest overall, demonstrating superior quality across a range of voice characteristics and languages. In Voice Arena, a human-judged evaluation platform, the model ranks highly in Japanese, Brazilian Portuguese, Vietnamese, and Modern Standard Arabic.

The ability to maintain audio quality and character consistency across extended recordings addresses a longstanding challenge in text-to-speech technology. Earlier models often struggled with voice drift or quality degradation when generating long passages, but Gemini 3.8 Flash TTS can sustain natural-sounding speech for hours while preserving emotional nuance and speaker identity.

Where Can Creators Access These Models?

Both Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are available through Google AI Studio, making them accessible to independent creators, small studios, and enterprises looking to integrate advanced speech synthesis into their workflows. The models support over 100 languages and dialects, expanding opportunities for creators working in non-English markets and underrepresented language communities.

The release of these models reflects a broader trend in Google DeepMind's AI strategy, which has included other generative tools like Veo for video creation and Lyria for music generation. By offering multiple specialized models for different creative tasks, Google is positioning itself as a comprehensive platform for AI-assisted content production.