We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through Multiple Choice Question Answering (MCQA) and open-ended long-form tasks. To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech. We evaluate model behavior across both real-time speech-to-speech and speech-to-text architectures. Our results reveal that voice-conditioned bias is a model-specific property. Furthermore, we demonstrate that personalization instructions consistently exacerbate demographic disparities. Our findings establish that voice bias is a controllable model characteristic, providing a foundational framework for future bias mitigation and evaluation in Audio-LLM development.
Figures & tables
Figure 1 : Demographic distribution of HEAR benchmark
Family
Model
System Prompt
Gender
Age
Language
DMD(Gender)
DMD(Age)
VCD(Female,Male)
VCD(Old,Young)
VCD(Non-EN, EN)
S2S
GPT-Realtime
Default
7.38%
-12.73%
18.42%
15.66%
5.39%
Personalized
9.48%
-15.15%
23.94%
18.21%
9.28%
Gemini Realtime
Default
5.12%
-8.25%
12.57%
1.03%
11.31%
Personalized
11.09%
-16.03%
26.05%
-3.15%
16.55%
S2T
Qwen-Omni
Default
10.78%
-17.00%
26.34%
-7.43%
9.37%
Table 1 : Spoken BBQ - Voice-Based Bias Metrics. VCD and DMD metrics comparison across different system prompts. The distribution of metrics are statistically different based on Mann-Whitney U test.
Model
Negative Question
Non-Negative Question
SIR
Average FR
SIR
Average FR
S2S
GPT Realtime
12.62%
10.70%
17.31%
11.24%
Gemini Realtime
3.62%
6.39%
6.05%
6.73%
S2T
Qwen-Omni
12.58%
17.19%
27.01%
16.11%
Gemma 4
8.68%
19.02%
16.69%
14.19%
Table 2 : Spoken BBQ - Content-based bias metrics
Figure 2 : Content-Level Bias Distribution across different question topic categories with negative framed questions.
Default SP
Personalized SP
Model
Disparity
Helpful
Disparity
Helpful
S2S
GPT-Realtime
8.00%
86.15%
6.00%
73.91%
Gemini Realtime
18.00%
32.62%
6.00%
31.74%
S2T
Qwen-Omni
22.00%
76.85%
54.00%
68.76%
Gemma4
10.00%
50.10%
8.00%
47.66%
Table 3 : Open QA - Voice bias metrics under different system prompts.
Figure 3 : Examples of Qwen-Omni response disparity across gender groups.
Graduate Institute of Communication Engineering National Taiwan University Taipei, Taiwan · NVIDIA Research NVIDIA Taipei, Taiwan · AI Center of Research Excellence National Taiwan University Taipei, Taiwan