CourionAI
EN
Newsletter
← All news
google 3 min read

Google's New Voice Models Keep Talking While They Think

Google shipped Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech to speech models built for voice assistants. The headline trick is that the slower one reasons and speaks at the same time instead of going silent.

Overlapping speech bubbles merging into a continuous ribbon with cogwheels turning inside

Google released two new voice models yesterday. Both are what the industry calls speech to speech models, meaning they listen to audio and answer with audio directly, without first turning your words into text and then turning their answer back into sound. Cutting out those two steps is what makes a spoken conversation feel like a conversation rather than a walkie talkie exchange.

The cheaper one, Gemini 3.8 Live, is aimed at things like support lines and in app assistants where cost per minute matters. The bigger one, Gemini 3.8 Live Extended Thinking, is built for requests that take several steps. Google reports that Extended Thinking took the top spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6, scored 68.6 percent on the tau-Voice test of whether a voice agent actually completes a task, 35.1 percent on Sierra’s harder banking version of that test, and 97.7 percent on Big Bench Audio. Artificial Analysis is an independent evaluator, which is worth more than a vendor’s own chart, though Google chose which numbers to print. The smaller 3.8 Live came second in the Speech Agent Arena, where people compare answers and vote.

Two features stand out beyond the scores. The models detect and switch between 97 languages in the middle of a conversation, so you can start in German and drift into English without touching a setting. And they run tools in the background while still speaking to you, using filler like “Let me check that” and then narrating progress, rather than going quiet for eight seconds while something loads. All generated audio carries SynthID, Google’s inaudible watermark for marking AI output.

What is behind this

Silence is the thing that breaks voice assistants. A text chatbot can pause for five seconds and nobody minds, because you can see it working. In speech, five seconds of nothing reads as a dropped call, so assistants have historically been forced to stay shallow and fast. Splitting the release in two is Google admitting there is no single right answer: one model for cheap, high volume chatter, one that thinks harder and covers the delay by talking through it. That is less a research breakthrough than an honest bit of product design, and it is why the “reasons and speaks simultaneously” line matters more than any benchmark on the page.

What this means for you: If you are curious rather than building anything, the shortest path is Search Live, where the cheaper model is available to everyone now. Gemini Live plus Docs, Gmail and Keep get the smarter model, though the Workspace parts need a Google AI Pro or Ultra subscription. Expect more customer service lines to sound unnervingly natural over the next year, and keep in mind that a fluent voice is not evidence of a correct answer. If you build things, both models are live in the Gemini API and Google AI Studio today, and the enterprise version is still private preview.

Sources

Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/

Next story

Google Is Making Laptops Again, and This Time the Selling Point Is the Assistant

Googlebook is a new laptop category built by Acer, Asus, Dell, HP and Lenovo around Google's Gemini Intelligence. Pre-orders open on 21 September in the US, with devices shipping this autumn.

An open laptop on a museum plinth with a glowing geometric form rising out of its screen