Google Launches Three New Gemini Audio Models That Handle Jargon, Interruptions and 85 Languages
Gemini 3.5 Transcribe, 3.5 Live and 3.5 Live Experimental are rolling out now, and the transcription model is a first for the Gemini family.

Key points
- Google released three new Gemini Audio models on 22 July 2025: Gemini 3.5 Transcribe, 3.5 Live and 3.5 Live Experimental.
- Gemini 3.5 Transcribe supports more than 85 languages and can tell apart up to three different speakers in a pre-recorded audio file.
- The new models handle background noise and mid-sentence interruptions better than their predecessors, Google says.
- Gemini 3.5 Transcribe is available today in English for macOS Gemini app users and on Android via the Rambler dictation feature in select countries.
- Developers can access all three models in public preview through the Gemini API.
Google has quietly shipped three new voice and transcription models under the Gemini Audio label, and at least one of them is something the company has never offered before.
The biggest newcomer is Gemini 3.5 Transcribe, a dedicated speech-to-text model, meaning software that listens to spoken audio and converts it into written text. Google says it replaces an older model called Chirp 3, with particular improvements to multilingual accuracy and fewer wording errors. Those are meaningful claims for anyone who has ever watched a transcription tool mangle a medical term or a proper noun.
What can it actually do?
Quite a lot, for a transcription tool. It strips out filler words like "um" and "uh" automatically, formats the text for you, and lets you supply a custom vocabulary list so that specialist jargon or unusual spellings are recognised from the start rather than corrected away.
That last feature matters for doctors, lawyers, engineers or anyone whose work lives in a specialised vocabulary. Instead of fixing errors by hand after every recording, you give the model your word list once.
The model can also identify up to three separate speakers in a pre-recorded audio file and attach word-level timestamps, so you know exactly who said what and when. First reported by The Verge, these details were confirmed in Google's own product announcement.
Who gets access right now?
Rollout is limited but real. English-language users on the macOS Gemini app get it today, as do Android users in select countries through a dictation feature called Rambler. Developers can try all three models through the Gemini API, the programming interface that lets software teams build Google's AI into their own products, via a testing environment called AI Studio. Google says Chrome browser support is coming soon.
| Model | What it does | Status |
|---|---|---|
| Gemini 3.5 Transcribe | Converts speech to text; strips filler words; tags speakers | Available now (English, macOS + Android Rambler) |
| Gemini 3.5 Live | Real-time voice chat; better with interruptions and accents | Available now |
| Gemini 3.5 Live Experimental | Narrates its own reasoning aloud during complex tasks | Available now (developer preview) |
The other two models, Gemini 3.5 Live and Gemini 3.5 Live Experimental, update the voice chat technology already inside Gemini. The standard Live model copes better when you cut yourself off mid-sentence or switch languages. The Experimental version goes further: it talks through its own thinking in real time as it works on harder questions, which can make it easier to catch when it goes wrong.
One thing worth noting
Google promised to release a Gemini 3.5 Pro model in June. It has not appeared yet. These audio models are separate from that and do not fill that gap. Worth keeping in mind before assuming Google's roadmap is back on schedule.
The honest takeaway: if you record meetings, interviews or medical notes and spend time fixing transcription errors by hand, Gemini 3.5 Transcribe is worth testing the moment it reaches your device. Set up a custom vocabulary list on day one. That single step will save more time than any other feature on offer here.



