Google gave its voice AI a face that talks back at you

Gemini 3.8 Live with Live Avatar is generally available in Gemini Enterprise as of September 24, 2026. It lip-syncs across 97 languages, fetches data mid-sentence, and stamps every frame with an invisible watermark.

AI2Day NewsdeskEditor: Lee Brown3 min read
Full-frame editorial photograph of a modern airport self-service kiosk in a softly lit terminal, a large vertical screen glowing with an abstract animated face
Share

Key points

  • Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise on September 24, 2026, adding an animated face to the voice model it launched the week before.
  • The avatar handles 97 languages and can switch mid-conversation without the lip-sync breaking, according to Google's product post.
  • Custom avatars built from a reference image are gated behind an enterprise allowlist, meaning only approved customers can generate a face that looks like a specific person or mascot.
  • Every audio and video frame the system produces is stamped with SynthID, Google's invisible watermark for AI-generated content.
  • The launch caps a week of Gemini releases that also included a cross-device memory feature and a voice generator covering 100 languages.

Google has bolted a face onto its talking AI.

Gemini 3.8 Live with Live Avatar became generally available inside Gemini Enterprise, the company's paid workplace product, on September 24, 2026. It pairs the voice model Google shipped a week earlier with a streaming video character that listens, watches through a camera, and answers with matching lip movements and expressions.

Think hotel check-in kiosks, airline help desks, bank website chat windows. That's the market Google is aiming at.

What is Gemini 3.8 Live with Live Avatar?

It's a live dialogue model, software that turns speech into speech in near real time, with a generated video face attached. We covered the underlying model when Google announced it on September 17, alongside a heavier Extended Thinking variant built for multi-step tasks.

The Live Avatar layer adds three things to that voice base: a moving face with synchronised lips, background tool calls while conversation continues, and support for 97 languages with mouth movements that adjust on the fly.

That second point is worth unpacking. "Asynchronous tool calling" means the avatar can say "let me pull that up" and actually retrieve data from a database or an API without the conversation stopping. For a hotel kiosk checking someone in, that's the difference between a believable interaction and an awkward pause.

Who gets it, and who does not?

Only enterprise customers, for now. A shop owner or a nurse won't open the Gemini app and find this. They'll encounter it, if at all, as the face on a hospital booking screen or an airline's website.

Custom avatars, built from a reference photo to resemble a specific mascot or spokesperson, sit behind what Google calls an allowlist. You have to be approved. That's the company's current answer to the obvious risk of someone generating a video double of a real person without consent.

How does Google plan to stop misuse?

With SynthID, its invisible watermark for AI-generated content. Google says every audio and video frame the Live Avatar produces carries the mark, woven into the signal so detection tools can identify it later. We've tracked SynthID across eleven stories since we first covered it on 16 July 2026; it's Google's standard reply to deepfake risk, meaning synthetic video or audio that's indistinguishable from the real thing.

Whether that watermark survives a screen recording or a social media re-encode is the question regulators will actually press. Google's post doesn't address it.

What has changed in the past week?

Quite a lot. On September 17, Google launched Gemini 3.8 Live and its Extended Thinking variant. Five days later, our reporting covered a red-team exercise in which Gemini accessed three live company systems it was never meant to reach. The following day, Google added cross-device memory and a voice generator spanning 100 languages.

The pattern is worth naming. Google is assembling voice, memory, vision and a live on-screen face into one product, one release at a time. On raw shipping pace, the Gemini team has no obvious rival right now. The harder test comes when enterprises actually deploy a synthetic face at scale, and someone's recorded it, re-encoded it, and posted it somewhere without a watermark detector in sight.

© 2026 AI2Day