Google Introduces Gemini 3.8 Live Audio Models for Real-Time Voice AI Workflows

Google DeepMind's Gemini Audio lineup now includes Gemini 3.8 Live for near real-time voice interfaces and Extended Thinking for more complex reasoning tasks.

Google Introduces Gemini 3.8 Live Audio Models for Real-Time Voice AI Workflows
Gemini 3.8 Live: Voice AI Models for Natural Dialogue

Google DeepMind has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two audio-focused models designed for more natural, near real-time conversations with AI. The models expand the Gemini Audio family beyond transcription, translation, and text-to-speech capabilities, with one aimed at high-volume voice interactions and the other built for more demanding reasoning while a conversation continues.

For businesses, the development matters because voice can become a more practical interface for AI-supported work. Rather than requiring users to type every instruction, a voice system could support conversational task handling, customer-facing interactions, or guided internal workflows. The important distinction is that the two models serve different levels of complexity, and Google has not published explicit public pricing for either model on its product page.

Google describes the models on its official Gemini Audio page, which also identifies access routes through Google AI Studio, the Gemini Live API, the Gemini API, the Gemini app, and related Google services.

What Gemini 3.8 Live and Extended Thinking are designed to do

Gemini 3.8 Live is positioned for near real-time voice interfaces. Google describes it as providing conversational capabilities optimized for high-volume, cost-effective use, alongside near-real-time reasoning. That positioning makes it the more direct fit for applications where responsiveness and the ability to handle many interactions are central requirements.

Gemini 3.8 Live Extended Thinking is the higher-end option. Google says it is intended for complex reasoning and can orchestrate multiple agents to solve background tasks while maintaining a natural conversation. In practical terms, this points to voice interactions that do more than answer a straightforward spoken request. A system could continue speaking naturally with a user while coordinating a more involved task in the background.

Model Google's stated focus Relevant workflow implication
Gemini 3.8 Live Near real-time voice interfaces, conversational capabilities, high-volume and cost-effective use Potential fit for responsive voice interactions at scale
Gemini 3.8 Live Extended Thinking Complex reasoning and orchestration of multiple agents for background tasks during conversation Potential fit for voice-led workflows that require more involved task coordination

The table reflects Google's stated positioning, not a benchmark comparison. The official material does not provide performance measurements, token costs, latency figures, or a public price list that would allow a more detailed operational comparison.

How the models fit into Gemini Audio

The new live models sit within a broader Gemini Audio portfolio. Google's product page also highlights Gemini 3.5 Transcribe, Gemini 3.5 Live Translate, and Gemini 3.1 Flash TTS. Those capabilities address distinct audio tasks: turning speech into text, translating live audio, and generating speech.

Gemini 3.8 Live and Extended Thinking represent a different emphasis. Their focus is the conversational layer itself, especially real-time dialogue and reasoning. This matters for teams assessing voice AI because a transcription or text-to-speech feature alone is not the same as a system designed to hold a live exchange and act on the context of that exchange.

Google also highlights SynthID watermarking across its audio work. SynthID is intended to flag whether speech has been AI-generated or edited. For organizations considering generated voice in customer or employee interactions, that safety capability is relevant, although the product page does not set out a complete implementation process or policy framework for individual use cases.

Access, pricing, and implementation questions

Google identifies multiple paths to the new capabilities. Developers can explore them through Google AI Studio and build through the Gemini Live API and Gemini API. Consumer access is also referenced through the Gemini app and related Google services. These routes indicate that Gemini Audio is intended to span experimentation, development, and end-user experiences rather than remain limited to a single product surface.

What remains less clear is the commercial and technical detail a business would need before committing to a production deployment. The official page does not list explicit public pricing for Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking. It also does not provide the operational details needed to determine which model will be more economical for a particular workload.

Before building around a voice model, teams should establish:

  • Whether the workflow needs fast conversational responses or more complex background reasoning.
  • Which access path fits the intended product or internal process, such as AI Studio experimentation or an API integration.
  • How voice input, generated speech, and task outputs will connect to existing business systems.
  • What testing is needed to assess the quality of responses for the organization's real conversations and tasks.

The multi-agent capability associated with Extended Thinking is particularly notable, but it should not be read as a ready-made business process. The value will depend on how reliably an implementation connects the model to the specific tools, data, and steps that make up the underlying work.

Voice AI creates an opportunity to reduce friction in tasks that begin with a spoken request, but a useful deployment still requires a clear handoff between conversation and action. That could mean routing information to an existing application, triggering an approved workflow, or keeping a user involved where a decision needs review.

Voice interfaces are only valuable when they connect reliably to the work that follows the conversation. Scalevise helps businesses turn promising AI capabilities into practical processes, from mapping suitable voice-led tasks to connecting models with the tools teams already use. Our AI workflow automation service can help reduce manual handoffs and build workflows with clear operational value. Discuss an AI automation project with Scalevise.

Frequently Asked Questions

What is Gemini 3.8 Live?

Gemini 3.8 Live is a Gemini Audio model for near real-time voice interfaces. Google describes it as offering conversational capabilities optimized for high-volume, cost-effective use and near-real-time reasoning.

What is Gemini 3.8 Live Extended Thinking?

Gemini 3.8 Live Extended Thinking is a Gemini Audio model aimed at more complex reasoning. Google says it can orchestrate multiple agents to solve background tasks while maintaining a natural conversation.

How can developers access Gemini 3.8 Live models?

Google identifies Google AI Studio, the Gemini Live API, and the Gemini API as developer access paths. The Gemini app and related Google services are also listed as consumer access routes.

Has Google published pricing for Gemini 3.8 Live and Extended Thinking?

The official Gemini Audio page does not include explicit public pricing for Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking.

What safety feature does Google highlight for Gemini Audio?

Google highlights SynthID watermarking, which is designed to flag whether speech was AI-generated or edited.


Conclusion

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking extend Google's audio strategy from individual speech tasks toward real-time, conversational AI. The division between high-volume live interaction and more complex reasoning gives developers a clearer starting point for evaluating voice-led workflows. Access is available through Google's developer and consumer channels, while pricing and deployment-specific performance details remain matters to confirm through the relevant Google portals and testing.