Audio LLMs Know When They Can't Hear You
Organizations: Apple · University of California San Diego
Abstract
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provide limited signals for detecting transcription failures. In contrast, we discover that transcription reliability is strongly represented in the model's audio-encoder representations. Based on this observation, we devise a lightweight reliability predictor that operates on representations extracted by the frozen audio encoder and predicts the reliability class before generation. The reliability predictor can trigger a clarification request from the user when their voice query is predicted to be unreliable, while allowing reliable queries to proceed without modifying the underlying Audio LLM. Our predictor achieves 81.10% in-domain and 78.09% cross-domain macro-F1 scores, outperforming the strongest baselines by 10.33 and 11.93 points, respectively. Finally, we show that reliability labels can transfer across Audio LLM families, and that transfer performance is closely related to the alignment of their model-specific reliability boundaries.
Figures & tables
| In-domain Test | Cross-domain Test | |||||
| Method | Macro-F1 | Accuracy | MAE | Macro-F1 | Accuracy | MAE |
| Acoustic corruption proxy | ||||||
| SNR | 57.93 | 65.14 | 0.39 | 51.59 | 60.85 | 0.47 |
| No-reference speech quality and intelligibility predictors | ||||||
| Audiobox-Aesthetics (PQ) ( Tjandra et al., 2025 ) | 22.43 | 31.30 | 1.10 | 25.03 | 34.26 | 0.91 |
| DNSMOS (OVRL) ( Reddy et al., 2021 ) | 39.74 | 47.88 | 0.60 | 41.76 | 51.34 | 0.59 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Clean Fail. | Reverb Fail. | SNR Gap | Order Viol. | Retained |
|---|---|---|---|---|---|
| Qwen2-Audio-7B-Instruct | 3.8% | 1.6% | 7.3% | 4.0% | 83.3% |
| Phi-4-Multimodal-Instruct | 2.1% | 1.0% | 3.8% | 5.1% | 88.1% |
| MOSS-Audio-8B | 8.5% | 4.0% | 3.4% | 4.4% | 79.7% |
| In-domain Test | Cross-domain Test | |||||
| Method | Macro-F1 | Accuracy | MAE | Macro-F1 | Accuracy | MAE |
| Acoustic corruption proxy | ||||||
| SNR | 60.72 | 68.38 | 0.36 | 54.39 | 64.09 | 0.43 |
| No-reference speech quality and intelligibility predictors | ||||||
| Audiobox-Aesthetics (PQ) ( Tjandra et al., 2025 ) | 23.01 | 32.88 | 1.12 | 27.38 | 36.70 | 0.94 |
| DNSMOS (OVRL) ( Reddy et al., 2021 ) | 38.14 | 45.43 | 0.64 | 41.05 | 49.57 | 0.64 |
| In-domain Test | Cross-domain Test | |||||
| Method | Macro-F1 | Accuracy | MAE | Macro-F1 | Accuracy | MAE |
| Acoustic corruption proxy | ||||||
| SNR | 55.43 | 67.77 | 0.36 | 51.99 | 64.00 | 0.42 |
| No-reference speech quality and intelligibility predictors | ||||||
| Audiobox-Aesthetics (PQ) ( Tjandra et al., 2025 ) | 22.90 | 28.97 | 1.05 | 22.05 | 28.26 | 0.87 |
| DNSMOS (OVRL) ( Reddy et al., 2021 ) | 38.70 | 47.17 | 0.59 | 42.17 | 52.08 | 0.57 |
| Parameters | FLOPs | |||||
|---|---|---|---|---|---|---|
| Model | Encoder | Decoder | Head | Encoder | Decoder | Head |
| Qwen2-Audio-7B-Instruct | 636.97M | 7.76B | 0.33M | 1.89T | 4.21T | 2.12M |
| Phi-4-Multimodal-Instruct | 441.24M | 3.84B | 0.26M | 114.20G | 1.32T | 1.70M |
| MOSS-Audio-8B | 643.62M | 8.19B | 329.99K | 243.05G | 2.67T | 2.12M |
| Pooling Method | Macro-F1 | Accuracy | MAE |
|---|---|---|---|
| Max Pooling | 77.60 | 79.50 | 0.22 |
| Mean Pooling | 79.69 | 81.35 | 0.20 |
| Learned Temporal Pooling | 81.10 | 82.60 | 0.18 |
| Transfer setting | Direct Transfer | Aligned Transfer | |||
|---|---|---|---|---|---|
| Alignment method | Global Mean | Global Median | Per-Boundary Mean | Per-Boundary Median | |
| MOSS Qwen2 | 9.69 | 1.90 | 1.92 | 2.30 | 1.70 |
| MOSS Phi-4 | 12.86 | 2.71 | 2.29 | 2.21 | 2.05 |
| Mean | 11.28 | 2.31 | 2.10 | 2.26 | 1.88 |