Google Releases Gemini 3.8 Live for Real-Time Voice Apps
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking for developers on September 15. Available through the Live API in the Gemini API and Google AI Studio, the models are intended for applications that maintain real-time conversations while reasoning and carrying out tasks.
Both models support asynchronous function calling, allowing an agent to make API and tool calls in the background while continuing to stream audio responses. Google also lists live visual context, precise handling of alphanumeric information, incremental merging of audio with structured data, and coverage for more than 97 languages.
Extended Thinking adds a configurable reasoning mode for complex, multi-step requests. According to Google, it can process such work in the background while responding to the user or narrating its progress in the main conversation. Google says the model ranks first on Artificial Analysis’ Speech-to-Speech leaderboard, but its announcement provides no testing details or independent reproduction results.
Google estimates pricing for Gemini 3.8 Live and Extended Thinking at $0.005 per minute of audio input and $0.018 per minute of audio output. A footnote says this estimate is based on rates of $3 per million input tokens and $12 per million output tokens. Access is also offered through integration partners Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel and Vision Agents.
The audio lineup also includes Gemini 3.5 Transcribe, released a month earlier for converting streamed speech into text. Google reports an average word error rate of 4.0% in streaming mode and 2.6% in non-streaming use. The model supports more than 85 languages, automatically handles language switching and accepts a custom vocabulary list containing up to 1,000 terms.
Smart Transcription mode produces structured, reader-ready text, accounts for speakers’ self-corrections and removes filler words. Through the Interactions API, the model can transcribe audio files up to one hour long with structured timestamps and speaker labels. The error-rate figures and broader performance claims come from Google; the announcement does not identify the evaluation datasets or methodology.
Practical context: In practical terms, the lineup provides developers with two distinct building blocks: Gemini 3.8 Live for two-way voice agents that use actions and context, and Gemini 3.5 Transcribe for captioning, call analysis and other text-only tasks. Asynchronous function calls are particularly relevant when external services are involved because the conversation need not stop while a request runs, although actual latency will also depend on the connected infrastructure.
| Parameter | Value |
|---|---|
| Access channel | Live API in the Gemini API and Google AI Studio |
| Audio input | $0.005 per minute |
| Audio output | $0.018 per minute |
| Integrations | Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, Vision Agents |
Sources
Event date: 2026-09-15. Primary source date: 2026-09-15.