VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation
Organizations: Veris AI
Abstract
Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests. The agent has sixteen tools backed by a stateful Stripe billing twin and the Apache Fineract loan engine, with account access blocked until caller verification succeeds. The tasks use public household electricity data and a policy based on Pennsylvania's residential billing rules. An LLM-as-a-verifier checks the agent's actions and spoken figures against explicit requirements. On a calibration run, it agrees with a code verifier on 99.1% of checks. Across fourteen voice stacks and three repeats per task, completion ranges from 17.3% to 44.7%. Grok Voice leads, and Gemini 3.8 Live and GPT-Live follow at about the same cost per call. Background television reduces pooled completion from 38.7% to 8.6%. The simulated caller often accepts an incorrect result because it hears the agent's words but cannot inspect its actions. These findings show why voice agents need evaluation across the whole call, including what they say, what they change, and how they handle competing speech.
Figures & tables
| VAmoS Bench | VAmoS Energy | |
|---|---|---|
| Backend | Seeded PostgreSQL | Stripe twin + Apache Fineract |
| Agent tools | 5 | 16 |
| Verification | Card digits, name, address/phone | Account number, name, second factor |
| Requests per call | One, sometimes multi-step | 2–4 |
| Data | Generated | ResStock usage; Pennsylvania policy |
| Caller personas | Per-scenario style | Calm, angry |
| Stack | Complete (%) | 95% interval | Latency (s) | Cost ($) | |
|---|---|---|---|---|---|
| Grok Voice | 44.7 | [38.8, 50.7] | 264 | 2.64 | 0.210 |
| Gemini 3.8 Live | 40.1 | [34.6, 45.9] | 284 | 1.66 | 0.193 |
| OpenAI GPT-Live 1 | 39.2 | [33.5, 45.3] | 260 | 1.62 | 0.192 |
| OpenAI Realtime 2.1 | 37.9 | [32.3, 43.8] | 269 | 2.12 | 0.633 |
| ElevenLabs | 36.4 | [30.9, 42.2] | 275 | 2.18 | 0.272 |
| OpenAI Realtime | 33.6 | [28.2, 39.5] | 265 | 2.32 | 0.607 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Check | Reads | Checks | terra agrees | luna agrees |
|---|---|---|---|---|
| Exact world effects, in order | tool results | 97 | 96 | 96 |
| No account tool before verification | tool calls | 97 | 97 | 97 |
| Required tool called | tool calls | 20 | 20 | 20 |
| Forbidden tool not called | tool calls | 34 | 34 | 34 |
| Quote, caller turn, then create | tool calls + turns | 34 | 32 | 33 |
| No secret spoken before verification | speech | 97 | 95 | 95 |
| Requests | Persona | Use case | |||||
| Stack | 2 | 3 | 4 | Calm | Angry | High bill | Payment |
| Grok Voice | 81 | 46 | 40 | 39 | 50 | 40 | 48 |
| Gemini 3.8 Live | 56 | 49 | 34 | 34 | 46 | 31 | 45 |
| OpenAI GPT-Live 1 | 65 | 45 | 34 | 44 | 34 | 38 | 40 |
| OpenAI Realtime 2.1 | 67 | 51 | 28 | 34 | 41 | 26 | 45 |
| ElevenLabs | 71 | 38 | 32 | 31 | 41 | 35 | 37 |
| Stack | All (%) | Rank | Counted (%) | Rank | Excluded | Caller | Verifier | Lost | Passes | Soundex |
|---|---|---|---|---|---|---|---|---|---|---|
| Grok Voice | 41.0 | 1 | 44.7 | 1 | 36 | 25 | 11 | 0 | 5 | 43 |
| Gemini 3.8 Live | 38.0 | 2 | 40.1 | 2 | 16 | 11 | 2 | 3 | 0 | 25 |
| OpenAI GPT-Live 1 | 37.3 | 3 | 39.2 | 3 | 40 | 30 | 10 | 0 | 10 | 26 |
| OpenAI Realtime 2.1 | 35.0 | 4 | 37.9 | 4 | 31 | 23 | 6 | 2 | 3 | 28 |
| ElevenLabs | 34.3 | 5 | 36.4 | 5 | 25 | 13 | 4 | 8 | 3 | 28 |
| OpenAI Realtime | 30.3 | 6 | 33.6 | 6 | 35 | 31 | 4 | 0 | 2 | 25 |
| Stack | Class | Models (as run) |
|---|---|---|
| Grok Voice | Speech-to-speech | grok-voice-think-fast-2.0 , voice eve |
| Gemini 3.8 Live | Speech-to-speech | gemini-3.8-live , same bridge as Gemini 3.1 Live |
| OpenAI GPT-Live 1 | Speech-to-speech | gpt-live-1 , tool use delegated to gpt-5.6-terra , voice marin |
| OpenAI Realtime 2.1 | Speech-to-speech | gpt-realtime-2.1 , same bridge as OpenAI Realtime |
| ElevenLabs | Bundled platform | Conversational AI: ElevenLabs ASR gpt-4.1-mini eleven_flash_v2 |
| OpenAI Realtime | Speech-to-speech | gpt-realtime-2 , voice alloy , server VAD (threshold 0.5) |
| Request | Tasks |
|---|---|
| Enable paperless after confirming the email | 73 |
| Report the current balance and delinquency status | 50 |
| Reject a one-month arrangement | 25 |
| Record the caller’s exact partial instalment | 25 |
| Record a payment made today | 25 |
| Quote before creating an arrangement | 22 |