Google's Gemini can now invent a voice from scratch, in 100 languages

Two new text-to-speech models let anyone design a custom character voice, dub in dozens of accents, or clone their own voice from a 30-second sample.

AI2Day NewsdeskEditor: Lee Brown4 min read
A photoreal editorial still of a modern recording studio microphone glowing with soft blue and violet light, wrapped in faint translucent waveforms suggesting s
Share

Key points

  • Google's new Gemini 3.8 Flash TTS and Flash-Lite TTS models can generate voices from a plain-English prompt across more than 100 languages and dialects.
  • Voice cloning needs only a 30-second sample plus a spoken consent recording from the voice's owner before the system will build it.
  • Every clip is stamped with SynthID, an inaudible watermark Google uses to mark AI-generated audio so it can be detected later.
  • Gemini 3.8 Flash TTS took first place on Hume AI's Voice Design Benchmark with a score of 71.4, and both models rank first and second on Hume AI's Overall Quality Index.
  • Developers can try both models today inside Google AI Studio, with partners including Figma, HeyGen and Wondercraft already plugging them in.

Google has just turned its Gemini assistant into something closer to a voice acting studio.

On Tuesday the company released two new text-to-speech models, software that turns written words into spoken audio. They're called Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, and they can invent a voice from a written description, dub a script across dozens of languages, or clone a real person's voice from a short sample.

Type "gravelly narrator with a soft Scottish lilt, slightly out of breath" and the model hands you a voice that fits. Paste in a script, mark up where you want a sigh or a gasp, and it performs them.

What can these models actually do?

They generate speech with the kind of direction you'd give a voice actor. The Flash version is built for creative work: audiobooks, game characters, podcasts. Flash-Lite is the cheaper workhorse, designed for high-volume jobs like dubbing a video library or running a call centre bot.

Google says the library ships with more than 2,000 ready-made voices, covering regional flavours like Mexican Spanish and Quebec French. You can build a brand new one and save it so your podcast host sounds consistent week to week.

The line-by-line direction is the part that stands out. Write stage cues and the model performs them. Non-verbal cues like <laughs> or <gasp> are supported directly in the script, alongside active-listening interjections for natural conversational texture.

How does the voice cloning work, and is it safe?

Cloning needs a 30-second audio sample, plus a spoken consent recording from the person whose voice it is. The system checks the consent clip against the reference before it'll build the voice.

Every clip the model produces carries SynthID, Google DeepMind's inaudible watermark for AI-generated audio. We've covered SynthID across ten stories since first reporting on it on 16 July 2026, but this is the first time it's been applied to a text-to-speech product aimed squarely at everyday creators.

That's a real safeguard, not a perfect one. Watermarks can degrade if audio is heavily re-encoded, and consent checks only stop the honest user. Someone determined to fake your boss's voice isn't filing paperwork with Google; they're using a shadier tool. Still, for the mainstream creators this product is aimed at, the guardrails are meaningfully tighter than a year ago.

How this fits with the rest of Gemini's audio push

Google's been shipping audio models at a steady clip. Earlier this month it released Gemini Live models that can think and talk simultaneously, and the family already includes Gemini 3.5 Live Translate for near real-time speech-to-speech translation and Gemini 3.5 Transcribe for speech-to-text. The new TTS models fill the last obvious gap: Google can now listen, translate, speak and clone, all inside one product family.

Model Best for Notable claim
Gemini 3.8 Flash TTS Character voices, audiobooks, games No. 1 on Hume AI Voice Design Benchmark (71.4)
Gemini 3.8 Flash-Lite TTS Bulk dubbing, voice agents No. 2 on Hume AI Overall Quality Index
Gemini 3.5 Live Translate Live speech-to-speech translation 70+ languages
Gemini 3.5 Transcribe Speech-to-text Handles jargon and background noise

The interesting fight here isn't the benchmark score. It's distribution. Google can drop these voices straight into Notebook and Vids, where hundreds of millions of people already are. That's a harder moat to cross than a leaderboard position.

What should ordinary users do?

If you make podcasts, videos or training material, it's worth a look inside Google AI Studio today. Pricing for the Flash-Lite tier is aimed at high-volume use, which typically lands well below hiring voice talent for routine jobs. Google hasn't published per-character rates for these specific models yet.

If you don't build things, the practical point is simpler. Assume the voice on the other end of a strange call might not be a person. If your bank or your boss sounds slightly off and asks for money or a code, hang up and call back on a number you already know.

© 2026 AI2Day