FAST TAKE · 2026-09-15 · GEMINI 3.8 LIVE + EXTENDED THINKING
Gemini 3.8 Live: real-time voice, plus an Extended Thinking sibling that keeps reasoning while you talk
Google made Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking generally available on September 15, 2026. Both are hosted real-time models with text, image, audio and video input, audio output and a 131K input window. Extended Thinking can keep reasoning and calling tools in the background while the audio conversation continues.
Verdict: Google adds a low-latency voice model and a background-reasoning sibling; useful hosted infrastructure, not a local-model event
The take
The standard Live model is built for low-latency audio-to-audio interaction, interleaved reasoning and asynchronous function calls. Google says it can switch automatically across 97 languages. Extended Thinking separates the conversational stream from longer background reasoning/tool work, so a voice agent does not have to go silent while it completes a task.
Google's Standard-tier rate card lists text input at $0.75 per million tokens, audio input at $3, text output at $4.50 and audio output at $12; the page also exposes minute equivalents for media. Those mixed modalities do not fit our text-only 50/50 calculator without creating a misleading blended figure, so we publish the prices here instead of adding a calculator row.
Our call: meaningful for hosted voice-agent builders, irrelevant to the local hardware ladder. Qwen3-Omni, Step-Audio and modular local ASR/TTS remain the privacy-first paths; Gemini Live wins when managed latency, multilingual reach and tool orchestration matter more than offline operation.
Where this fits
Models: Qwen3-Omni-30B-A3B-Instruct · Step-Audio 2 mini · MiniCPM-o 4.5 · Chatterbox (Turbo + Multilingual v3)
Hardware: NVIDIA RTX 5090 · M5 Max MacBook Pro 64 GB
Sources
Next step
Try this in the planner→