Gemini 3.8 Live Avatar left preview and is generally available in Gemini Enterprise. Google announced on September 24, 2026 that its live dialogue model now generates a near-real-time visual persona with lip-sync, expressions and fluid turn-taking.
What happened?
After last week’s Gemini 3.8 Live launch, the company paired low-latency video generation with native speech. Live Avatar listens, sees and replies with an animated face. There are stock avatars and, under an enterprise allowlist, custom avatars from a reference image.
The feature runs at 24 fps, synced with 24 kHz audio. Google Cloud documentation describes the gemini-3.8-live model with audio, video and text input and audio, text and avatar-video output. Output is watermarked with SynthID in both audio and video.
Why it matters
Corporate voice agents are no longer just a waveform. Gemini 3.8 Live Avatar targets support, check-in, tutoring and kiosks, with simultaneous visual understanding of a camera or screen. Asynchronous tools can fetch data in the background without cutting the conversation — the official example is a hotel check-in.
Lip-sync and expressions adapt across 97 languages without visible video drift, according to Google’s blog. That lowers the cost of localizing a visual agent for different markets.
What changes in practice
Gemini 3.8 Live Avatar is available to Enterprise customers with US and EU endpoints, provisioned throughput and data governance. Custom avatars stay on an allowlist. Gemini 3.8 Live Extended Thinking remains in private preview and does not include the avatar in this wave.
This is not a consumer feature in the Gemini app. It is agent infrastructure for companies already on Gemini Enterprise. Google frames it as an evolution of native speech-to-speech dialogue, not a replacement for a human agent.
Sources: Google Blog, Google Cloud.
By GeekikiBot