cs.CLSep 27, 2026

NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models

Authors: Ziwei Chen

Organizations: University of California San Diego

Abstract

We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final turn. Our pilot contains 22 human-verified pairs (44 audio conditions) and evaluates five models on NSV perception, emotion understanding, and response adaptation. Results show that models generally perform better at detecting NSVs than at interpreting their fine-grained emotional meaning or producing appropriately differentiated responses. The data construction pipeline, dataset, and evaluation pipeline are publicly available at https://github.com/ChenzwNina/nsv-construction.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. NVV-SuperBench: Beyond Words, Beyond Quality-Benchmarking Nonverbal Vocalizations in Speech Generation

    Apr 17, 2026Liumeng Xue, Weizhen Bian, Jiahao Pan +9Seed-Tts-Eval Benchmark

  2. NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

    Jun 14, 2026Jialong Mai, Jinxin Ji, Xiaofen Xing +2Seed-Tts-Eval BenchmarkVocalizations

  3. Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

    Jun 19, 2026Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou +4Automatic Speaker VerificationSpeaker