Blog

Gemini 3.8 Live and Extended Thinking are now on Glytos

16 September 2026 · Glytos

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September. A day later both were running on Glytos. No waitlist and no private beta: open a voice agent, pick one, and place a call.

Two models, two jobs

Both are speech-to-speech models. One model hears the caller and answers in speech, instead of a chain of separate services passing the conversation along. Where they differ is what they are for.

Gemini 3.8 Live is the one Google builds for scale and cost efficiency, and describes as the default for most low-latency voice agents. Google says it detects and moves between 97 supported languages in the middle of a conversation, and it comes with thirty voices to choose from.

Gemini 3.8 Live Extended Thinking is built for harder work: tasks with several steps, where the answer has to be reasoned out rather than recalled. Google describes it as reasoning and speaking at the same time, so the conversation does not stop while it thinks. In Google's words, it acknowledges a request early with something like "Let me check that…" and narrates a longer task while it is still running.

They join GPT-Live on Glytos, so a voice agent can now run on the speech model from either OpenAI or Google.

What Google's numbers show

Every figure in this section is Google's. Artificial Analysis and Sierra ran the benchmarks with each model at its high thinking setting, and Google published the results with its evaluation methodology. We have not reproduced any of them.

On Artificial Analysis' Speech to Speech Index, a broad measure of speech-to-speech quality, Extended Thinking takes the top spot at 82.6%. Gemini 3.8 Live scores 76.0%.

Artificial Analysis Speech to Speech Index: Gemini 3.8 Live Extended Thinking 82.6%, GPT-Live-1 Astra 81.5%, Grok Voice Think Fast 2.0 81.3%, Gemini 3.8 Live 76.0%, Gemini 3.1 Flash Live 71.5% at high and 63.9% at minimal thinking

Chart from Google's announcement, data by Artificial Analysis.

The picture changes on agentic tasks, where the model has to carry a job through to the end rather than hold a good conversation. On Artificial Analysis' τ-Voice, Extended Thinking completes 68.6% of tasks and Gemini 3.8 Live 30.1%.

Artificial Analysis agentic performance on τ-Voice: Gemini 3.8 Live Extended Thinking 68.6%, GPT-Live-1 Astra 67.9%, Grok Voice Think Fast 2.0 56.5%, Gemini 3.1 Flash Live at high thinking 37.7%, Gemini 3.8 Live 30.1%, Gemini 3.1 Flash Live at minimal thinking 26.2%

Chart from Google's announcement, data by Artificial Analysis.

Sierra's τ³-Banking is harder still. The agent has to find its way through a large, unstructured knowledge base and make several tool calls in a row to resolve a realistic banking request. Extended Thinking leads there at 35.1%.

Sierra τ³-Banking leaderboard: Gemini 3.8 Live Extended Thinking 35.1%, GPT-Live-1 Astra 32.0%, xAI-Realtime 16.5%, Gemini 3.1 Flash Live 11.3%, GPT-Realtime 2 10.3%

Chart from Google's announcement, data by Sierra.

How we read them

A few things stand out, and none of them needs more than the charts to see.

Thinking pays off where a call has steps. Between the two Gemini models, the gap on overall quality is 6.6 points. On agentic tasks it is 38.5, and Extended Thinking more than doubles the score. On that chart Gemini 3.8 Live even sits below the previous generation run at its high thinking setting. These are not a good model and a better one. They are two models built for different jobs.

The lead at the top is narrow. On both Artificial Analysis charts, the next model is within about a point of Extended Thinking, and on τ³-Banking the gap is about three. That is a place in the leading group, and a point on a benchmark is not something a caller hears.

Multi-step voice work is not solved. The best score on τ³-Banking completes about a third of the tasks. Whatever model an agent runs on, a flow where a wrong step is expensive deserves careful testing before it goes live.

Fluency and accuracy pull against each other. In ServiceNow's EVA-Bench, published in the same methodology document, Gemini 3.8 Live scores highest of all on conversational experience, while Extended Thinking gives up some of that experience in exchange for accuracy. That fits Google's advice to make 3.8 Live the default.

Which one to pick

Start on Gemini 3.8 Live. An agent that greets callers, answers questions, qualifies a lead or books into a single tool is doing the job it was built for, and Google builds it to answer without reasoning-induced delays.

Move to Extended Thinking when the call itself has to think: several tool calls in sequence, a policy with exceptions, or an answer put together from more than one place in a knowledge base. Thinking takes work, so use it where a call needs it rather than everywhere.

The model is a setting on the agent, so trying both on the same agent means changing one field, not rebuilding anything.

What it is like on a real call

We called both, and both held a natural conversation.

On the short, simple calls we placed, the two sounded much alike. That is what Google's own numbers would predict: the difference they measure is in multi-step tasks, not in small talk.

That is an impression from a handful of calls rather than a benchmark. We have not measured either model, and we would rather say nothing than publish a number we have not taken.

Trying it

On a voice agent, open the model settings, set the speech engine to Realtime, and choose Gemini Live as the provider. Then pick Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking as the model, choose a voice and call it. Everything else the agent already has keeps working: its prompt, its knowledge base, its tools, its phone number.

Early days

This is a first look rather than a verdict. The measurements that would turn an impression into a claim come later, and when they do they get their own post.

Gemini 3.8 Live and Extended Thinking are now on Glytos · Glytos