Google's voice model tops the index, finishes 35.1% of banking calls

Google shipped Gemini 3.8 Live and Extended Thinking on September 15, with Extended Thinking taking first place on Artificial Analysis' speech-to-speech index at 82.6. At the published per-minute rates, an hour of two-way audio runs about $1.38.

Google's voice model tops the index, finishes 35.1% of banking calls

On September 15, Google shipped two voice models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both are native speech-to-speech, both went live the same day in the Gemini Live API and AI Studio, and neither is available as open weights.

Extended Thinking took the top spot on Artificial Analysis’ Speech-to-Speech Quality Index with 82.6.

The leaderboard is crowded. Three other numbers are not.

Price the first-place finish first. Behind 82.6 sit GPT-Live-1 Astra (Medium) at 81.5 and Grok Voice Think Fast 2.0 (High) at 81.3. The top three are separated by 1.3 points. The ranking is news. The margin is not.

What matters is the same model against three different sets of tasks:

  • Big Bench Audio (audio reasoning) 97.7%
  • τ-Voice (general voice-agent task completion) 68.6%
  • Sierra’s τ-Voice-banking (banking phone tasks) 35.1%

One model, one release day, 97.7 down to 35.1.

Two more entries for the record: the non-reasoning Gemini 3.8 Live placed second in Speech Agent Arena, a human-preference evaluation, and on ServiceNow’s EVA-Bench Google reports the pair pushing the Pareto frontier for complex workflows.

Price stopped being the constraint a while ago

Published rates are $0.005/min for audio in and $0.018/min for audio out, which Google derives from $3/1M input tokens and $12/1M output tokens.

Run an hour of two-way conversation through that: $0.30 in, $1.08 out, roughly $1.38 total. That figure is backed out of Google’s own per-minute pricing and is not a number Google published.

$1.38 an hour. No staffed seat in any labor market lands near it. Whatever has kept voice agents out of the contact center for two years, cost was never it.

One line in the capability list names the actual reason: alphanumeric precision. Google specifically cites confirmation codes, claim numbers and technical data, and says outright that this is a common failure point in voice systems. Read that backwards and the last two years make sense. Voice agents were not failing to understand the customer. They were getting one digit of a claim number wrong.

The rest of the release reads the same way. Tool calls now run in the background while audio keeps streaming, so the caller is not parked in silence. The model switches among 97 languages mid-conversation and holds accent consistency. On long-running tasks it acknowledges out loud that it is going to check, then narrates progress while it works.

Where 35.1% lands on the org chart

Put the three scores side by side and the line through the customer-service job draws itself.

97.7% says the model understands and can reason. 68.6% says it finishes most generic, well-bounded voice tasks. 35.1% says something else entirely: on a call with an account, a compliance surface and real money attached, close to two-thirds of the time it does not get to the end.

We covered AI gutting customer service on July 30, Accenture and Google putting 1,000 forward-deployed engineers against YouTube support and cutting handle time 37% on September 8, and KB Insurance issuing employee numbers to six agents and slotting them into the grade ladder on September 11. All three point the same direction: whatever can be written into a script goes first. 35.1% puts a scale mark on that direction.

So tier-one handling takes the first hit. Balance checks, address changes, billing dates, a fixed script. That tier is already measured on handle time, has the shortest training cycle and the highest churn.

The second-order move is the one worth watching, because it is substitution rather than subtraction. The 65% of calls that do not complete get transferred to a person, and the transferred call is harder than the call used to be. The customer has already spent minutes with a machine and arrives annoyed, the context has to be reconstructed in the first three seconds, and the third of the task the machine did complete has to be verified rather than trusted. Tier-one seats shrink while the bar for the remaining seats rises. That is fine news for people already holding one and bad news for anyone who expected to break into the industry through that tier.

Watch the next revision of τ-Voice-banking. The day it moves from 35.1% to north of 50% is the day a bank can actually hand over its front line. Until then, any claim that AI has taken over the contact center can be checked against one number.


Sources

Keep reading

Andon Labs now sells the AI boss. Its own store has $60K left. AI & Jobs

Andon Labs now sells the AI boss. Its own store has $60K left.

Andon Labs opened Pion on September 14, a waitlisted platform for handing a business to persistent agents. The same week, a reporter walked into the company's own AI-run San Francisco shop and found no customers and $60,000 left of the original $100,000.

#andon-labs#agentic-ai#automation