Updates·September 16, 2026, 16:55

Google launches Gemini 3.8 Live with vision, voice and 97 languages

AI-generated and checked against the sources listed below.

Google has released Gemini 3.8 Live and 3.8 Live Extended Thinking, new voice models that can see, reason and solve tasks in the background while the conversation continues uninterrupted. The models support 97 languages and top a key benchmark for speech-to-speech quality.

AI-generated image

Google has released two new models in its Gemini Live series: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The models are built for real-time voice conversations and, according to Google, can handle interruptions, switch languages mid-conversation and carry out tasks in the background without the conversation stalling.

Sees images in real time

A key new feature is that the models can process live video in near real time and use it as context in the conversation. Among other things, Google shows examples where Gemini 3.8 Live guides an employee through onboarding based on what it sees, and follows along visually while playing chess.

The models automatically support 97 languages and can switch between them mid-conversation. At the same time, they can perform tool calls and API calls asynchronously while they keep talking. This means Gemini can say something like "let me just check that" and continue the conversation while a task is being solved in the background.

Top ranking on benchmarks

Gemini 3.8 Live Extended Thinking is the most advanced version and is made for deeper, multi-step reasoning while it talks. According to the sources, it ranks number one on Artificial Analysis' speech-to-speech quality index with a score of 82.6, ahead of GPT-Live-1 Astra, among others. It also scores 68.6 percent on the τ-Voice benchmark and 35.1 percent on Sierra's τ-Voice banking test for agent tasks, as well as 97.7 percent on Big Bench Audio. The regular Gemini 3.8 Live model ranks number two in Speech Agent Arena, and on ServiceNow's EVA-Bench, which tests voice agents on complex workflows, the models according to Google push the frontier for how well precision and natural conversation can be combined.

The models are already available to developers via the Gemini API and Google AI Studio, and to ordinary users via the Gemini app, Google Workspace and Search. For developers, there are also integrations via partners such as LiveKit, Pipecat, LangChain, Vercel, Agora and Fishjam, which handle the technical infrastructure for media streaming.

The price of the models isn't clearly stated in the available material. Google itself lists $0.005 per minute for audio input and $0.018 per minute for audio output, while another source mentions $0.84 per hour for standard Live audio input and $3.50 per hour for Extended Thinking (High). It isn't clear whether these are different methods of calculation.

New model for transcription

Alongside the launch, Gemini 3.5 Transcribe is also mentioned, a dedicated speech-to-text model released last month. It supports more than 85 languages and has an average word error rate (WER) of 4.0 percent in streaming and 2.6 percent without streaming. The model can be used for things like live captioning and call center agents, and via the Interactions API it can transcribe audio files of up to an hour with timestamps and speaker recognition.

For the ordinary AI user, the launch means that voice-controlled assistants are becoming more useful in practice for actual work, not just short questions and answers. That the model can watch along, understand several languages automatically and keep working on a task while you talk about something else points toward voice assistants that function as a real collaborator rather than just a command line with a microphone.

What it means for you

If you already use the Gemini app, Google Workspace or Search, you may soon experience a voice assistant that can watch along, understand you in several languages and solve tasks while you keep talking. If you're a developer, you can try the models now via the Gemini API and Google AI Studio, possibly through partners like LiveKit or LangChain. If you need to build speech-to-text solutions, for example captions or call center support, Gemini 3.5 Transcribe is worth a closer look.

Sources

More on this topic

Get the week's AI news in your inbox

Choose your level, topics and length. One email a week, unsubscribe at any time.

Subscribe to Promptly Newsletter
PromptlyNewsletterRSSLog in

The news on aijour is AI-generated and checked against the cited sources.