Now Reading
Google Launches Gemini 3.8 Live With Advanced Voice AI and Extended Thinking

Google Launches Gemini 3.8 Live With Advanced Voice AI and Extended Thinking

Google Gemini 3.8 Voice Models

Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026. The new live dialogue models target natural voice conversations, real-time interactions, and complex tasks that require deeper reasoning.

The company is rolling out the models across the Gemini API, Google AI Studio, Gemini Enterprise, Search Live, Gemini Live, and Google Workspace.

Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff, introduced the models on behalf of the Gemini Audio Team. Google describes them as its most advanced live dialogue models yet and says they offer stronger intelligence and parallel reasoning.

Moreover, Gemini 3.8 Live focuses on scale and cost efficiency while combining conversational intelligence, fluid dialogue, and visual grounding. Gemini 3.8 Live Extended Thinking, meanwhile, targets more complicated tasks that demand additional intelligence and multi-step reasoning.

According to Google, both models provide building blocks for reliable, production-ready voice agents. At the same time, they aim to make Gemini conversations across its app, Workspace, and Search more fluid.

Benchmark Results and Real-Time Capabilities

Google reports that Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis’ Speech to Speech Quality Index. It also reports scores of 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark for agentic task completion.

Additionally, Google reports a 97.7% result on Big Bench Audio. The company says Extended Thinking maintains a highly competitive price point compared with other frontier models.

Gemini 3.8 Live placed second in Artificial Analysis’ Speech Agent Arena, according to Google, which cited strong user preference. Furthermore, the company describes the model as cost-effective and suitable for deployment at scale.

On ServiceNow’s EVA-Bench, Google says its models push the Pareto Frontier for complex workflows. Specifically, it says they balance accuracy with conversational quality. The evaluation used the Live API on the Gemini Enterprise Agent Platform.

Beyond benchmark performance, Gemini 3.8 Live can process visual inputs in near real time. According to Google, this additional context helps the model provide more useful responses.

The model also detects and switches among 97 supported languages during conversations. Meanwhile, it can execute tools and API calls in the background without stopping the ongoing dialogue. It can acknowledge requests and continue speaking while those tasks finish.

For more demanding prompts, Extended Thinking can reason while speaking. Additionally, it uses early verbal cues to acknowledge prompts and provides live progress narration during multi-step background tasks. As a result, users can follow ongoing work without interrupting the conversational flow.

Demonstrations accompanying the announcement highlighted several possible applications. For example, Gemini 3.8 Live guided employee onboarding using visual context and answered questions in real time. Other demonstrations showed it playing chess, creating business plans and marketing toolkits, and providing step-by-step troubleshooting through Search Live.

Extended Thinking demonstrations focused on more complicated workflows. For instance, the model converted raw sketches and near-real-time voice feedback into functional React components. It also coordinated multi-step restaurant reservations through asynchronous function calls while maintaining the conversation.

Additionally, demonstrations showed Extended Thinking working across Docs Live, Gmail Live, and Keep Live in Google Workspace.

Live API, Developer Tools, and Safety

Several developer platforms use the Gemini Live API to support voice-driven applications, including Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms can manage real-time media streaming infrastructure while developers concentrate on their applications.

Google is also working with Salesforce, Genspark, and Lumeris on the models. In particular, the announcement highlighted interest in their latency, conversational fluidity, and tool-calling capabilities.

The Live API supports low-latency voice and vision interactions through continuous audio, image, and text streams. Moreover, it delivers spoken responses through a stateful WebSocket connection.

See Also
SoftBank Group logo on sign

The documented audio input uses raw 16-bit PCM at 16kHz, while audio output uses raw 16-bit PCM at 24kHz. The interface also accepts JPEG images at up to one frame per second, along with text.

Developers can choose between server-to-server and client-to-server implementations. In the first approach, a backend forwards client stream data to the Live API. Alternatively, frontend software can connect directly through WebSockets.

Potential applications cover a broad range of fields. They include retail assistants, customer support agents, interactive game characters, robotics, smart glasses, vehicles, education, translation, transcription, and captioning. Other documented applications include health companions and financial advisory tools.

Google also states that all audio generated by its AI products carries SynthID. The imperceptible watermark sits directly within generated audio so that systems can identify AI-created content and help counter misinformation.

Availability and Rollout

Both models began rolling out on September 15, 2026. Developers can access them through the Gemini API and Google AI Studio, while enterprises can access them through a private Gemini Enterprise preview.

Gemini 3.8 Live is also available to everyone through Search Live. Meanwhile, users can access Extended Thinking through Gemini Live.

Within Workspace, Google AI Pro and Ultra subscribers can use Extended Thinking in Docs. Additionally, all Google AI subscribers can access it through Gmail and Keep.

Google says Gemini Enterprise for Customer Experience will support both models soon. Furthermore, the company plans to bring Extended Thinking to Google Workspace business customers in a future rollout.

View Comments (0)

Leave a Reply

Your email address will not be published.

© 2024 The Technology Express. All Rights Reserved.