cs.SDSep 28, 2026

HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark for Audio Language Models

Authors: Shen Yan, Duc Le, Irina-Elena Veliche

Organizations: Meta Superintelligence Labs

Abstract

We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through Multiple Choice Question Answering (MCQA) and open-ended long-form tasks. To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech. We evaluate model behavior across both real-time speech-to-speech and speech-to-text architectures. Our results reveal that voice-conditioned bias is a model-specific property. Furthermore, we demonstrate that personalization instructions consistently exacerbate demographic disparities. Our findings establish that voice bias is a controllable model characteristic, providing a foundational framework for future bias mitigation and evaluation in Audio-LLM development.

Figures & tables

Explore similar work

CardsList
  1. VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

    Apr 19, 2026Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang +1Large Language Model BiasLarge Audio Language Models

  2. Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition

    Apr 23, 2026Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers +1Large Language Model DecodingAlgorithmic Fairness