cs.SDSep 2, 2026

VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis

Authors: Mengzhe Geng

Abstract

Speech systems increasingly infer how an utterance should be delivered from context, but a plausible delivery plan may not be supported by the input. VoxReason is a small public benchmark and verifier for testing this failure before waveform synthesis. Each of its 100 cases fixes the utterance, provides derived records that name the source emotion and intensity, and changes one licensed cue. A system must cite the record for its delivery decision and update only the plan fields associated with the edit. We use deterministic verifier references, not a model leaderboard. This holdout excludes every emotion and intensity combination observed in training. A prior-only predictor achieves 0.958 accuracy across plan fields but never changes its plan consistently after a cue edit. This contrast shows that plan agreement does not demonstrate source grounding. The released suite provides an auditable measurement layer for screening structured speech plans before synthesis. It is limited to derived records, not audio inputs.

Explore similar work

CardsList