Gemini 3.8 Live
Google's real-time speech-to-speech model for production voice agents
Google released Gemini 3.8 Live on September 15, 2026, alongside a reasoning-heavier sibling called Gemini 3.8 Live Extended Thinking. Both run over a WebSocket connection through the Gemini Live API, take audio, video, images, and text as input, and reply in audio, understanding and speaking 97 languages with the ability to switch languages mid-conversation. Pricing runs $0.005 per minute of audio input and $0.018 per minute of audio output, which works out to roughly $1.38 for an hour of conversation.
The models can run a function or query an API in the background while still talking, drop in natural filler like “let me check that” while they work, and support near real-time visual grounding. Google is positioning them against OpenAI’s GPT Live 1 Astra and xAI’s Grok Voice, and is also making them available as an enterprise private preview inside Gemini Enterprise and Search Live.