Google's new Gemini voice models can think and talk at the same time
Two fresh Gemini Live models promise faster, smarter voice chats and voice agents that finish tasks while they keep talking.

Key points
- Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-first AI models, rolling out from today.
- Extended Thinking scored 82.6 on Artificial Analysis' Speech to Speech Quality Index, taking the top overall spot.
- The lighter 3.8 Live model handles 97 languages and switches between them mid-conversation.
- Both models run tools in the background while continuing to chat, with verbal cues like "Let me check that..."
- Consumers get the models in Search Live, the Gemini app, and Google Docs and Gmail for paid subscribers.
Google has pushed out two new versions of Gemini, its family of AI models, built specifically for talking out loud.
The pair are called Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Think of them as Google's assistant getting a sharper brain and a smoother mouth together. The lighter one is tuned for speed and cost. The heavier Extended Thinking version is built for tricky, multi-step jobs where the assistant has to reason before it answers.
My read: the interesting trick here isn't the benchmark scores, it's that these things keep chatting while they work. That is what will make voice assistants feel less like walkie-talkies and more like a colleague on speakerphone.
What is actually new?
Both models are designed for near real-time voice conversation, and both can run tools, meaning calling other software or looking things up, in the background while the conversation carries on.
Ask the assistant to check your calendar and book a table, and it says "Let me check that..." out loud, keeps chatting, and narrates progress as the task finishes. No awkward silence while a spinner turns.
Gemini 3.8 Live also handles visual input in near real-time. Point a camera at a broken appliance and it talks you through what it sees. It supports 97 languages and switches between them mid-sentence without you changing a setting.
Extended Thinking reasons and speaks at the same time, so you get early verbal cues instead of dead air while it works out a longer answer.
When Google DeepMind released the 3.5 generation of audio models in late August, which we covered on 26 August, the headline feature was better handling of jargon and interruptions. This update goes further: the reasoning happens during the conversation, not before it.
How good is it, really?
Google DeepMind is leaning hard on independent benchmarks. Extended Thinking took first place on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, and posted 97.7% on Big Bench Audio, a test of audio reasoning.
On agent tasks, where the AI has to complete a job rather than just chat, it scored 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark. The lighter 3.8 Live came second in the Speech Agent Arena, a user preference ranking.
| Model | Best benchmark result |
|---|---|
| Gemini 3.8 Live Extended Thinking | 82.6 on Speech to Speech Quality Index |
| Gemini 3.8 Live Extended Thinking | 97.7% on Big Bench Audio |
| Gemini 3.8 Live Extended Thinking | 68.6% on τ-Voice agent tasks |
| Gemini 3.8 Live | 2nd in Speech Agent Arena |
Topping the speech quality chart is a real result. The agent scores are harder to read without knowing how other frontier models sit on the same scale.
Where can normal people use it?
Both models start rolling out today. Everyone gets 3.8 Live inside Search Live, Google's talk-to-search feature, for real-time troubleshooting help.
Extended Thinking is going into Gemini Live, the voice mode of the Gemini app. Paying Google AI Pro and Ultra subscribers get it inside Google Docs; all Google AI subscribers get it in Gmail and Keep.
Developers can pull both from the Gemini API and Google AI Studio. Businesses get them in private preview inside Gemini Enterprise, with partners including Salesforce and Lumeris already testing.
One detail worth flagging for anyone worried about fake audio: Google says every clip these models generate carries SynthID, an invisible watermark baked into the sound so AI-made audio can be detected later.
Common questions
Do I need to pay to try it?
No. The lighter Gemini 3.8 Live is rolling into Search Live for everyone. Extended Thinking is in the free Gemini app's voice mode, with extras for paid subscribers in Docs, Gmail and Keep.
Will it work in my language?
Probably. Gemini 3.8 Live supports 97 languages and can switch between them mid-conversation without you having to change a setting.



