
Gemini 3.8 Live and Extended Thinking: Production-grade voice agents with real-time reasoning and tool-calling
Published by AINave Editorial • Reviewed by Ramit
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models yet, designed for production-grade voice agents that reason and execute tools without breaking conversational flow. Both models are available today via the Gemini Live API and Google AI Studio, and they target a specific pain point: replacing cascaded speech pipelines that chain ASR, an LLM, and TTS with a single native speech-to-speech model.
Two models, two trade-offs
Gemini 3.8 Live is built for scale and cost efficiency. It combines conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking adds configurable multi-step reasoning that runs in the background while the model continues speaking, using early verbal cues like "Let me check that" and narrating progress as tasks execute. Google's demos show it converting sketches plus voice feedback into working React components and coordinating multi-step bookings.
Both models are hosted only. There is no self-hosted option. This matters for teams that need data residency or low-latency on-premise deployment.
Pricing that makes voice agents viable at scale
Pricing is $0.005 per minute for audio input and $0.018 per minute for audio output, based on tokenized cost estimates. At roughly $0.84 per hour for the base model, this undercuts most cascaded pipelines and puts continuous voice agents within reach for customer support, booking, and coding assistance use cases.
What the benchmarks actually show
Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis' Speech to Speech Quality Index, the highest reported score. It also achieved 68.6% on Tau-Voice and 35.1% on Sierra's Tau-Voice-banking benchmark for agentic task completion, and 97.7% on Big Bench Audio for audio reasoning. Gemini 3.8 Live placed second in the Speech Agent Arena, a human preference evaluation. On ServiceNow's EVA-Bench, Google reports the models push the Pareto Frontier for complex workflows, balancing task accuracy with conversational quality. These are vendor-reported results using third-party benchmarks, so independent verification is still valuable.
Integration partners and deployment options
Developers can integrate via the Live API through partners that handle real-time media streaming: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. Google also announced collaborations with Salesforce, Genspark, and Lumeris. Example apps are available on GitHub.
What builders should watch for
The main limitation is that these are hosted models with no self-hosted path. Pricing is based on token estimates, so actual costs may vary depending on turn-taking and audio length. The Extended Thinking model is more expensive per minute due to additional compute for reasoning. If you need strict data control or ultra-low latency, a self-hosted cascaded pipeline may still be necessary. Google's enterprise agent platform and tool-calling capabilities are the key differentiator here, but those require the Gemini Enterprise Agent Platform.
For teams building voice agents at scale, Gemini 3.8 Live offers a compelling alternative to chaining separate models. The decision comes down to whether the hosted trade-offs work for your use case.
FAQs
Sources
- Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
- Google News - Google releases Gemini 3.8 Live Extended Thinking...
- Google releases Gemini 3.8 Live and Extended Thinking for...
- Google Releases Gemini 3.8 Live for Production Grade Voice...
- Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking
- Gemini 3.8 Live powers Google Search Live
- Google’s new speech model Gemini 3.8 Live supports real-time reasoning
- Google Unveils Gemini 3.8 Live and Extended Thinking Models
- Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail, & Keep
- Google Launches Gemini 3.8 Live Models for Real-Time AI Dialogue
- Gemini 3.8 Live: Designing Voice Agents That Think Without Breaking the Conversation - DEV Community
- Google rolls out Gemini 3.8 Live and 3.8 Live Extended Thinking with parallel reasoning for voice AI
- Gemini 3.8 Live Review: Google Developer Breakdown for Business & AI Calling Leaders (2026) | Auto Interview AI
- Google Launches Gemini 3.8 Live and Extended Thinking Models





















