Google releases Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its “most expressive audio generation models yet”, with support for more than 100 languages
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet.
The release is also being positioned across Google’s developer and consumer surfaces, according to Google AI’s rollout announcement. That makes expressive speech generation a product capability rather than a standalone model demonstration.
First-order effects
Developers using Google AI Studio and the Gemini API gain Flash and Flash-Lite TTS options spanning more than 100 languages, expanding the set of markets they can address with Google’s speech stack.
Google’s Gemini and Google Vids surfaces gain a more expressive speech-generation layer, bringing the new models into end-user creation workflows as well as developer tools.
Second-order effects
Teams building multilingual voice experiences can consolidate more language coverage on Google’s TTS models, increasing pressure on other speech providers to match both expressive control and language breadth.
Google’s earlier speech-to-speech translation release becomes more useful alongside broader TTS coverage: translation and generated output can increasingly be assembled from the same audio-model family.
Third-order effects
If Google continues joining low-latency dialogue, translation and expressive output, voice products will compete less on basic transcription or playback and more on how naturally a single model stack handles a full conversation.
More capable voice design and replication make provenance measures more consequential for audio deployments; Google had already applied SynthID watermarking to its Flash Live audio model.
The trend: Google is building an integrated multilingual audio stack in which generation, real-time dialogue and translation reinforce one another across developer and consumer products.
introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, our most expressive audio generation models yet these models enable creators, developers, and enterprises to create richer, more expressive audio experiences try them via the Gemini API and in AI Studio: https://aist…
Introducing Gemini 3.8 Flash and Flash-Lite TTS, our new SOTA text to speech model with: - a new voice design experience - 2,000+ production ready voices - voice replication - support for 100 languages - voice remixing (soon) - #1 spot on Hume AI's voice benchmarks and more!!
3/ Gemini 3.8 Flash and Flash-Lite take the top two spots on Overall, at 0.92 and 0.91. Gemini 3.1 Flash, an older model, still leads Acting at 3.97, well clear of the 3.69 behind it.
Rolling out starting today: — Developers: Both models in @GoogleAIStudio and the Gemini API — Consumers: Gemini 3.8 Flash TTS in @Gemini_Notebook and Gemini 3.8 Flash-Lite TTS in Google Vids — Coming soon: Both models in Gemini Enterprise https://blog.google/...
finally, my AI voice clone can say “let's take this offline” and “you're on mute” in my exact voice on phone calls, while i build things in peace 😁 voice cloning has arrived in @googleaistudio, y'all!
I cant get over the fact that they explicitly added an active listening “mhm” feature. we really just built robots to pretend they care about what we are saying. Being able to direct it line by line like an actual voice actor is pretty crazy though.
Gemini 3.8 Flash TTS processes 44.1 characters per second, compared to 40.2 characters per second for Gemini 3.8 Flash-Lite TTS, approximately 2.7x and 2.4x faster than realtime, respectively. Both models remain behind faster Text to Speech models we track, such as Falcon 2 at 20…
Create and deploy custom audio with our new text-to-speech models: 🔵 Gemini 3.8 Flash TTS: Design unique voices with distinct accents and characteristics. 🔵 Gemini 3.8 Flash-Lite TTS: Built for efficiency and scale, choose from your created styles or our expansive production-read…
Gemini 3.8 Flash TTS is out, and it can sing (a bit) too. It works out the tune by itself. Style: Singing Transcript: Happy release day to you, Happy release day to you, Happy release day, dear Logan, Happy release day to you.
Proud of the team for delivering the most natural and expressive speech model I've ever used (by far). One simple insight was that text-to-speech isn't a “small model
Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Gemini 3.8 Flash TTS debuts at #1 on our Pronunciation Robustness Benchmark and #2 on our Provider Voice Arena Leaderboard Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are @GoogleDeepMind's latest Text …
Can you hear that? Our Gemini Audio family is getting louder 🔊 We're introducing two of our most expressive audio generation models yet from @GoogleDeepMind: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.
We're launching Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS ⚡️ Our most expressive audio models yet let you create custom voices across 100+ languages or pick from 2,000+ ready-to-use ones. You can direct back-and-forth conversations, guide the delivery line-by-line, and a…
Forget the leaderboard wins. The real headline in Google's Gemini 3.8 TTS launch is that cloning a voice from a sample lasting 30 seconds is now a standard developer feature. Google's safeguard is a matching spoken consent recording from the voice owner, plus SynthID watermarks a…
@Bangkok8ai Hello! You can definitely create and save your custom voice. You can use your own customer voice from an external source via our voice replication feature - you will be asked to verify that you have consent to use it.
I was invited to test Gemini 3.8 Flash TTS. This is currently my favorite model. The voice outputs from this model are extraordinary. I never really rated AI voice generation because of the uncanniness associated with low quality. This Gemini IMO crosses that chasm, like Opus 4.5…