The Most-Liked AI Voice Agent Fails 1 in 4 Real Tasks. It's Not Even Top 5 at Finishing Them.
On August 24, 2026, Artificial Analysis launched the Speech Agent Arena: paid, screened humans holding blind, live voice calls with two hidden AI agents on the same task, then voting on which one they preferred. Gemini 3.1 Flash Live Preview - Minimal wins that vote outright, 1,046 Elo, ahead of every other model tested. Its Task Success Rate on the exact same benchmark, whether it actually completed the booking or the order, is 74.6%, not close to the top 5. The model that finishes the most real tasks, SpaceXAI's Grok Voice Think Fast 2.0 at 94.7%, doesn't crack the top 5 for preference at all. Here's the full data trail on what a voice AI leaderboard is actually measuring, and why the two numbers you'd assume move together, don't.
TL;DR
- 📊 The preference leaderboard— Aug 24, 2026: Artificial Analysis's Speech Agent Arena has Gemini 3.1 Flash Live Preview - Minimal winning human preference at 1,046 Elo, ahead of Gemini 3.1 Flash Live Preview - High (1,014), GPT-Realtime-1.5 (1,000), GPT Realtime Aug '25 (944), and ElevenLabs Agents (937).
- ✅ The task-success leaderboard— on the same benchmark, Grok Voice Think Fast 2.0 High leads at 94.7% Task Success Rate. Gemini 3.1 Flash Live Preview - Minimal, the preference winner, isn't in that top 5 at all.
- ⚠️ The gap— that same Gemini Minimal scores 74.6% Task Success Rate on its own. Subtract from 100 and it fails roughly 1 in 4 real bookings, orders, or requests it's assigned.
- ⚡ Why preference wins anyway— Gemini Minimal answers in 0.96 seconds (Time to First Audio) versus 1.54 seconds for Qwen Audio 3.0 Realtime Plus. Preference tracks speed closely across this whole dataset, independent of whether the task got done.
- 💵 Price doesn't fix it— GPT-Realtime-2.1 High costs $10.75 per hour of input audio, 7.2x Gemini Minimal's $1.50, and still finishes fewer tasks (91.5%) than Grok Voice Think Fast 2.0 High at $4.80/hr (94.7%).
- 🧮 The fix that already exists— Artificial Analysis's own Speech to Speech Index weighs Speech Reasoning, Agentic Performance, Arena Preference, and Task Success Rate at 25% each. Almost none of the leaderboards actually quoted in launch announcements use that weighting.
What the Speech Agent Arena actually tests
Artificial Analysis has run text-model arenas for years, the same pairwise-preference, Elo-fitting format LMSYS popularized for chatbots. The Speech Agent Arena, announced August 24, 2026, does the same thing for voice: a paid, screened third-party participant is handed a scenario, has two separate live phone calls, one with each of two hidden Speech-to-Speech models, then records which conversation they preferred. Fifteen of the 35 scenarios are agentic, meaning the model has to actually do something through a tool call, book a dental appointment, order two pizzas and a side under a $45 budget, instead of just talking. Twenty are non-agentic conversation only.
The preference vote produces an Elo score, exactly like a chatbot arena. But because 15 of the scenarios require a real action, Artificial Analysis also runs a second, independent evaluation pipeline on those calls: did the model make the correct final tool call or calls, verified after the fact, with any participant deviation or unverifiable case excluded. That second number is Task Success Rate. It is not a proxy or a guess about whether the human sounded happy. It's whether the appointment got booked.
The leaderboard everyone will quote
Here's the headline number from the arena: which model do people prefer talking to, full stop.
Speech Agent Arena: Preference Elo, top 5
Source: Artificial Analysis, 'Announcing the Speech Agent Arena,' Aug 24, 2026
Google's Gemini 3.1 Flash Live, and specifically its cheapest “minimal” reasoning-effort setting, wins the popularity contest outright. Not the flagship “high” version of the same model. Not GPT-Realtime. Not ElevenLabs's cascaded stack. The small, fast, minimal-effort configuration is the one humans, blind to which model they're talking to, like the most.
The leaderboard almost nobody will quote
Now hold the same 35 scenarios and score them a different way: not who did people like talking to, but who actually finished the job.
Speech Agent Arena: Task Success Rate, top 5
Source: Artificial Analysis, 'Announcing the Speech Agent Arena,' Aug 24, 2026
Look at the name that's missing. Gemini 3.1 Flash Live Preview - Minimal, the model that just won the preference vote outright, isn't on this chart. It isn't close to being on this chart.
The specific number: 1,046 Elo, 74.6% Task Success
Artificial Analysis reports the figure directly: Gemini 3.1 Flash Live Preview - Minimal, the arena's most-preferred model at 1,046 Elo, has a Task Success Rate of 74.6%. Not 94. Not 90. Seventy-four point six.
Subtract 74.6 from 100 and you get 25.4. On the exact scenarios that require the model to book something, order something, or complete a tool call correctly, the single most-liked AI voice agent in this study fails roughly 1 out of every 4 times.
Artificial Analysis's own framing of the result makes the same point without editorializing: “a preferred conversation does not always result in successful task completion, some conversations can sound as though the requested action was completed even when the required final tool call was unsuccessful.” In plain terms: the model can sound like it booked your appointment and not have booked it.
Why the least reliable model wins the popularity vote
The arena data points at a clean mechanism: speed. Artificial Analysis measured Time-to-First-Audio (TTFA), how long a caller waits before the model starts talking back, against Preference Elo across the field, and the correlation is direct.
Time to First Audio vs. Preference Elo
Source: Artificial Analysis, 'Announcing the Speech Agent Arena,' Aug 24, 2026
Gemini Minimal responds in 0.96 seconds and tops the arena. Qwen Audio 3.0 Realtime Plus takes 1.54 seconds, barely half a second slower, and lands 347 Elo points lower. Artificial Analysis's own read: “highly preferred models tended to respond quickly, sound more natural and produce fewer unnatural sounds or audio artifacts.” None of that is about whether the model gets the job done. It's about whether the call feels good while it's happening.
Paying more doesn't buy you out of this either
The obvious assumption is that you get what you pay for, that the model priced highest also finishes the most tasks. The arena's pricing data says otherwise.
| Model | Price (per hour, input audio) | Task Success Rate |
|---|---|---|
| Gemini 3.1 Flash Live Preview - Minimal | $1.50 | 74.6% |
| Grok Voice Think Fast 2.0 High | $4.80 (3.2x Gemini) | 94.7% |
| GPT-Realtime-2.1 High | $10.75 (7.2x Gemini, 2.2x Grok Voice) | 91.5% |
Grok Voice Think Fast 2.0 High costs 3.2x what Gemini Minimal costs, and that premium buys a real 20-point jump in task success, a defensible trade if you actually need the task done. But GPT-Realtime-2.1 High costs more than double what Grok Voice costs, and 7.2x what Gemini costs, and it still finishes fewer tasks than the cheaper Grok Voice model. Price ordering and task-success ordering don't match anywhere on this table.
Lay all three numbers side by side
Put preference rank, task success, and price in one place and no model wins on all three.
| Model | Preference Elo rank | Task Success Rate | Price / hr |
|---|---|---|---|
| Gemini 3.1 Flash Live Preview - Minimal | #1 (1,046 Elo) | 74.6% | $1.50 |
| Grok Voice Think Fast 2.0 High | Outside top 5 | 94.7% (#1) | $4.80 |
| GPT-Realtime-2.1 High | Outside top 5 | 91.5% | $10.75 |
The model people like most isn't the model that gets the job done. The model that gets the job done isn't the cheapest option. And the most expensive model on this table isn't even the best at the one thing that expense is supposedly buying.
The index Artificial Analysis already built to fix this
Artificial Analysis isn't blind to the gap; its own Speech to Speech benchmarking methodology combines four components into one index, equal-weighted at 25% each: Speech Reasoning (Big Bench Audio), Agentic Performance (a benchmark it calls 𝜏-Voice), Arena Preference, and Task Success Rate. A model needs valid results across all four to even appear in the index. That's the right instinct: preference is one input among four, not the whole score.
Artificial Analysis Speech to Speech Index: component weighting
Source: Artificial Analysis Speech to Speech Benchmarking Methodology
The problem isn't that this index doesn't exist. It's that it isn't the number that travels. A launch tweet, a comparison thread, a “which voice AI should I use” post almost always quotes an arena Elo or a preference ranking alone, because that's the number that reads like a simple win. Task Success Rate requires reading past the headline. Most people don't.
Why this generalizes past voice agents
This exact failure mode isn't new to speech models. Chat arenas have shown the same split for years: a model can win an Elo-style preference vote for being agreeable, fast, and well-formatted, while scoring worse than a blunter, slower model on tasks with a verifiable right answer. Voice just makes the gap easier to see, because the mismatch is audible in real time: a warm, fast, confident voice that quietly didn't make the booking.
Judging a voice agent purely on preference is the same failure mode as judging a job candidate purely on how likeable they sounded in the interview. It's a real signal. It is not the same signal as whether the work got done.
If you're deploying any AI agent unsupervised, voice or otherwise, for customer calls, bookings, order-taking, the number that should gate the decision is Task Success Rate, or your own equivalent of it, measured on your actual workload. Preference is a useful tiebreaker between two models that both clear your success bar. It is a dangerous substitute for that bar.
Measure it yourself, don't inherit someone else's leaderboard
The deeper issue is that almost nobody running a voice agent in production is running Artificial Analysis's exact 35 scenarios. Your booking flow, your product catalog, your budget constraints are different from “order two pizzas and a side under $45.” A model's rank on someone else's arena tells you it can be preferred or reliable in general. It doesn't tell you it will be reliable on your workload, which is the only number that actually matters once real customers are calling in.
That's the same gap MegaBrainis built to close for every model call, not just voice. Routing through one API at zero markup means you get real cost and latency data per model, per call, so instead of trusting a vendor's Elo score or a leaderboard screenshot, you can run your own success-rate check against your own scenarios and see which model actually finishes your job, at what price, in production.
And if the agent you're building is meant to run the way a voice agent does, taking calls, completing tasks, unattended, around the clock, that's exactly the territory mhermes, MegaBrain's always-on agent runtime, is built for: a persistent, isolated VM per agent so it keeps working, keeps state, and keeps a record of what it actually completed, instead of you finding out days later that a “preferred” model quietly didn't finish the job.
Sign up at getmegabrain.com to route by what a model actually gets done, not by which one sounds best doing it.
MegaBrain Gateway
500+ models. One API. No markup.
Use in Claude Code, Cline, Cursor, or any coding agent.
Newsletter
Stay in the loop
Get the latest model comparisons and guides — no spam, unsubscribe anytime.