Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests. The agent has sixteen tools backed by a stateful Stripe billing twin and the Apache Fineract loan engine, with account access blocked until caller verification succeeds. The tasks use public household electricity data and a policy based on Pennsylvania's residential billing rules. An LLM-as-a-verifier checks the agent's actions and spoken figures against explicit requirements. On a calibration run, it agrees with a code verifier on 99.1% of checks. Across fourteen voice stacks and three repeats per task, completion ranges from 17.3% to 44.7%. Grok Voice leads, and Gemini 3.8 Live and GPT-Live follow at about the same cost per call. Background television reduces pooled completion from 38.7% to 8.6%. The simulated caller often accepts an incorrect result because it hears the agent's words but cannot inspect its actions. These findings show why voice agents need evaluation across the whole call, including what they say, what they change, and how they handle competing speech.
Figures & tables
VAmoS Bench
VAmoS Energy
Backend
Seeded PostgreSQL
Stripe twin + Apache Fineract
Agent tools
5
16
Verification
Card digits, name, address/phone
Account number, name, second factor
Requests per call
One, sometimes multi-step
2–4
Data
Generated
ResStock usage; Pennsylvania policy
Caller personas
Per-scenario style
Calm, angry
Table 1: The two VAmoS benchmarks. Observed completion depends on both the benchmark and the evaluated configurations.
Figure 1: Task construction and selection. The loop checks each added request against the account state left by earlier requests. Development runs inform the final selection; not every generated combination was run.
Stack
Complete (%)
95% interval
n
Latency (s)
Cost ($)
Grok Voice
44.7
[38.8, 50.7]
264
2.64
0.210
Gemini 3.8 Live
40.1
[34.6, 45.9]
284
1.66
0.193
OpenAI GPT-Live 1
39.2
[33.5, 45.3]
260
1.62
0.192
OpenAI Realtime 2.1
37.9
[32.3, 43.8]
269
2.12
0.633
ElevenLabs
36.4
[30.9, 42.2]
275
2.18
0.272
OpenAI Realtime
33.6
[28.2, 39.5]
265
2.32
0.607
Table 2: Completion, median response latency, and mean cost per call (Appendix D ). Intervals are Wilson 95%; n is the number of counted attempts. Latency and cost use calls with verdicts. Configurations are in Appendix F .
Figure 2: (a) Completion by background condition, with 46–75 counted attempts per stack and condition. (b) Transfers to a human by persona, among calls with verdicts and traces. All exam tasks are intended to be resolved by the agent.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Check
Reads
Checks
terra agrees
luna agrees
Exact world effects, in order
tool results
97
96
96
No account tool before verification
tool calls
97
97
97
Required tool called
tool calls
20
20
20
Forbidden tool not called
tool calls
34
34
34
Quote, caller turn, then create
tool calls + turns
34
32
33
No secret spoken before verification
speech
97
95
95
Appendix
Table 3: Agreement between the code verifier and two LLM verifiers, check by check, on 97 completed attempts from a calibration run of this exam. “Reads” is the evidence a check needs; the speech rows are the ones where the code extracts values with patterns.
Requests
Persona
Use case
Stack
2
3
4
Calm
Angry
High bill
Payment
Grok Voice
81
46
40
39
50
40
48
Gemini 3.8 Live
56
49
34
34
46
31
45
OpenAI GPT-Live 1
65
45
34
44
34
38
40
OpenAI Realtime 2.1
67
51
28
34
41
26
45
ElevenLabs
71
38
32
31
41
35
37
Appendix
Table 4: Completion (%) by number of requests, caller persona, and use case of the last request. Before exclusions, each stack has 18/90/192 attempts with 2/3/4 requests, 150 per persona, and 111/189 per use case; counted denominators are smaller.
Figure 3: Completion against median response latency and mean cost per call. Vertical bars are Wilson 95% intervals; the cost axis is logarithmic.
Stack
All (%)
Rank
Counted (%)
Rank
Excluded
Caller
Verifier
Lost
Passes
Soundex
Grok Voice
41.0
1
44.7
1
36
25
11
0
5
43
Gemini 3.8 Live
38.0
2
40.1
2
16
11
2
3
0
25
OpenAI GPT-Live 1
37.3
3
39.2
3
40
30
10
0
10
26
OpenAI Realtime 2.1
35.0
4
37.9
4
31
23
6
2
3
28
ElevenLabs
34.3
5
36.4
5
25
13
4
8
3
28
OpenAI Realtime
30.3
6
33.6
6
35
31
4
0
2
25
Appendix
Table 5: Completion over all attempts and over counted attempts, with the attempts excluded and why: the simulated caller broke its script (“caller”), the verifier misjudged the call (“verifier”), or the platform lost the session (“lost”). “Passes” is the number of excluded attempts that had passed. “Soundex” is the number of passes, over all attempts, whose caller was verified only through the sound-alike name match (described below the table). Ordered by counted completion.
Stack
Class
Models (as run)
Grok Voice
Speech-to-speech
grok-voice-think-fast-2.0 , voice eve
Gemini 3.8 Live
Speech-to-speech
gemini-3.8-live , same bridge as Gemini 3.1 Live
OpenAI GPT-Live 1
Speech-to-speech
gpt-live-1 , tool use delegated to gpt-5.6-terra , voice marin
OpenAI Realtime 2.1
Speech-to-speech
gpt-realtime-2.1 , same bridge as OpenAI Realtime
ElevenLabs
Bundled platform
Conversational AI: ElevenLabs ASR → gpt-4.1-mini → eleven_flash_v2
OpenAI Realtime
Speech-to-speech
gpt-realtime-2 , voice alloy , server VAD (threshold 0.5)
Appendix
Table 6: The fourteen stacks in this run, in order of observed completion. Every stack runs the same prompt, tool declarations, and dispatcher from the shared core. Models are as recorded in the candidate logs and the implementations’ configuration.
Request
Tasks
Enable paperless after confirming the email
73
Report the current balance and delinquency status
50
Reject a one-month arrangement
25
Record the caller’s exact partial instalment
25
Record a payment made today
25
Quote before creating an arrangement
22
Appendix
Table 7: The requests the 100 tasks draw on, with the number of tasks that include each.
Production voice agents span cascaded, speech-to-speech, and hybrid architectures. Voice-agent benchmarks typically measure component quality and conversational properties such as word error rate, latency, naturalness, and turn-taking. Fewer measure whether the agent handled a phone call correctly on its own. Contact centers refer to this as ``containment'': the share of phone calls the automated system resolves without handing off to a human. On some phone calls the right outcome is refusal or a redirect. To address this gap, we introduce VAmoS Bench, the Voice Agent Simulation Bench. It measures complete voice-agent systems end to end in a stateful customer-support task. The agent is Riley, a credit-card support representative for a fictional bank who can freeze, cancel, replace, or activate a card. Each of 100 scenarios supplies a simulated caller with a private goal and a seeded PostgreSQL backend. The platform uses each scenario to populate and activate an isolated simulation in which the caller reaches Riley over audio; roughly one-third apply adversarial pressure. The agent can use five tools that execute real SQL against the backend. Each scenario also defines binary assertions. A grader evaluates them against the complete trace of what the caller and agent said and what the agent did, including tool invocations, arguments, and returned rows. This catches an agent that claims to have changed a card without updating the database, as well as one that makes the right database change while disclosing protected information. This first benchmark version focuses on financial services. Its evaluation protocol supports an evolving leaderboard: additional voice agents can be evaluated on the same version, while later versions can expand the tasks and scenarios.
Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool calling evaluations without re-annotating the tool schema and gold labels. Our dataset-agnostic framework uses text-to-speech, speaker variation, and environmental noise to create paired text-audio instances while preserving the original dataset annotations. Based on extensive evaluation of 7 omni-modal models on audio-converted versions of Confetti and When2Call, our framework demonstrates that the performance is strongly model- and task-dependent: Gemini-3.1-Flash-Live obtains the highest Confetti score (70.4), whereas GPT-Realtime-1.5 performs best on When2Call (71.9). On Confetti, the text-to-voice gap ranges from 1.8 points for Qwen3-Omni to 4.8 points for GPT-Realtime-1.5. A targeted analysis of failure cases demonstrates that degradations most often reflect misunderstandings of argument values in the speech. Considering real-world deployment scenarios, we further report text-only results, an ambiguity-based reformulation stress test, and a reference-free LLM-as-judge protocol validated against human preferences. Notably, we find that open-source Qwen3 judges with at least 8B parameters exceed 80% agreement with proprietary judges, supporting privacy-preserving evaluation. Overall, our framework provides a verifiable and reproducible first-stage diagnostic that complements purpose-built audio corpora.
Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversation simulation and comprehensive voice-specific evaluation. We present EVA-Bench, an end-to-end evaluation framework that addresses both. On the simulation side, EVA-Bench orchestrates dynamic bot-to-bot audio conversations with automatic simulation validation that detects user simulator error and appropriately regenerates conversations before scoring. On the measurement side, EVA-Bench introduces two composite metrics: EVA-A (Accuracy) and EVA-X (Experience). EVA-Bench includes 213 scenarios across three enterprise domains, a controlled perturbation suite for accent and noise robustness, and multi-trial measurements that distinguish peak from reliable capability. Across 12 systems spanning all three architectures, we find: (1) no system simultaneously exceeds 0.5 on both EVA-A pass@1 and EVA-X pass@1; (2) peak and reliable performance diverge substantially (median pass@k--pass^k gap of 0.44 on EVA-A); and (3) accent and noise perturbations expose substantial robustness gaps, with effects varying across architectures, systems, and metrics (mean Δ up to 0.314). We release EVA-Bench under an open-source license.