Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provide limited signals for detecting transcription failures. In contrast, we discover that transcription reliability is strongly represented in the model's audio-encoder representations. Based on this observation, we devise a lightweight reliability predictor that operates on representations extracted by the frozen audio encoder and predicts the reliability class before generation. The reliability predictor can trigger a clarification request from the user when their voice query is predicted to be unreliable, while allowing reliable queries to proceed without modifying the underlying Audio LLM. Our predictor achieves 81.10% in-domain and 78.09% cross-domain macro-F1 scores, outperforming the strongest baselines by 10.33 and 11.93 points, respectively. Finally, we show that reliability labels can transfer across Audio LLM families, and that transfer performance is closely related to the alignment of their model-specific reliability boundaries.
Figures & tables
Figure 1: Audio LLMs are poor judges of their own transcription reliability. Bars show the probability of predicting reliable transcription within realized WER bins for acoustically corrupted speech. Both zero-shot and two-shot self-assessments are poorly aligned with actual reliability, with only modest improvement from two-shot prompting.
Figure 2: Critical SNR varies substantially across speech–corruption pairs. Distributions show the critical SNR required for Qwen2-Audio-7B-Instruct to achieve a WER no greater than 10% under background noise, noise with reverberation, and music. The variation within each corruption category shows that a fixed SNR threshold cannot reliably characterize transcription reliability. Dashed lines indicate category means.
Figure 3: Overview of the reliability dataset construction pipeline. Starting from clean speech and a fixed acoustic corruption, we vary the corruption strength and obtain a transcription from the target Audio LLM. The generated transcript is compared with the reference transcript to compute WER. We then search for the corruption levels at which the model crosses the WER boundaries defining our reliability classes and store the resulting pair-specific boundaries in a label bank.
Figure 4: Audio-encoder-based transcription reliability prediction. The pretrained audio encoder is kept frozen and produces a sequence of frame-level representations. A learned temporal pooling module assigns a scalar weight to each frame and aggregates the weighted representations into a single utterance embedding. A lightweight classifier then predicts one of four reliability classes: reliable, minor degradation, moderate degradation, or severe degradation.
In-domain Test
Cross-domain Test
Method
Macro-F1 ↑
Accuracy ↑
MAE ↓
Macro-F1 ↑
Accuracy ↑
MAE ↓
Acoustic corruption proxy
SNR
57.93
65.14
0.39
51.59
60.85
0.47
No-reference speech quality and intelligibility predictors
Audiobox-Aesthetics (PQ) ( Tjandra et al., 2025 )
22.43
31.30
1.10
25.03
34.26
0.91
DNSMOS (OVRL) ( Reddy et al., 2021 )
39.74
47.88
0.60
41.76
51.34
0.59
Table 1: Four-class transcription reliability prediction and cross-domain generalization for Qwen2-Audio-7B-Instruct. For scalar baselines, three ordinal decision thresholds are calibrated using the training split and fixed during evaluation. Macro-F1, accuracy, and mean absolute error (MAE) are reported on the in-domain and cross-domain test sets.
Figure 5: Cross-model transfer of reliability label banks. The left panel reports transfer accuracy when training a reliability predictor using label banks constructed by different Audio LLMs. The middle panel shows the relationship between critical-SNR mismatch and transfer degradation: blue circles show direct cross-model transfer, with larger mismatch associated with greater degradation, while yellow triangles show transfer after aligning the source critical-SNR boundaries using boundary-specific median shifts estimated from only 5% of the shared speech–corruption pairs. This alignment substantially reduces the transfer gap for the two transfers to MOSS. The right panel evaluates the estimation of critical-SNR mismatch from different fractions of the shared pairs, reporting the mean and standard deviation over 20 random subsets. The estimate approaches the full-label-bank value even when using only a small fraction of the shared pairs.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Clean Fail.
Reverb Fail.
SNR Gap
Order Viol.
Retained
Qwen2-Audio-7B-Instruct
3.8%
1.6%
7.3%
4.0%
83.3%
Phi-4-Multimodal-Instruct
2.1%
1.0%
3.8%
5.1%
88.1%
MOSS-Audio-8B
8.5%
4.0%
3.4%
4.4%
79.7%
Appendix
Table 2: Effect of quality-control filtering on candidate speech–corruption pairs. We report the percentage of candidate pairs removed by each filtering criterion for each target Audio LLM, together with the fraction retained for label-bank construction.
In-domain Test
Cross-domain Test
Method
Macro-F1 ↑
Accuracy ↑
MAE ↓
Macro-F1 ↑
Accuracy ↑
MAE ↓
Acoustic corruption proxy
SNR
60.72
68.38
0.36
54.39
64.09
0.43
No-reference speech quality and intelligibility predictors
Audiobox-Aesthetics (PQ) ( Tjandra et al., 2025 )
23.01
32.88
1.12
27.38
36.70
0.94
DNSMOS (OVRL) ( Reddy et al., 2021 )
38.14
45.43
0.64
41.05
49.57
0.64
Appendix
Table 3: Four-class transcription reliability prediction and cross-domain generalization for Phi-4-Multimodal-Instruct. We report macro-F1, accuracy, and mean absolute error (MAE) on the in-domain and cross-domain test sets, following the same baselines and evaluation protocol as the Qwen2-Audio-7B-Instruct results in Table 1 . The proposed reliability predictor achieves the strongest overall performance in both settings.
In-domain Test
Cross-domain Test
Method
Macro-F1 ↑
Accuracy ↑
MAE ↓
Macro-F1 ↑
Accuracy ↑
MAE ↓
Acoustic corruption proxy
SNR
55.43
67.77
0.36
51.99
64.00
0.42
No-reference speech quality and intelligibility predictors
Audiobox-Aesthetics (PQ) ( Tjandra et al., 2025 )
22.90
28.97
1.05
22.05
28.26
0.87
DNSMOS (OVRL) ( Reddy et al., 2021 )
38.70
47.17
0.59
42.17
52.08
0.57
Appendix
Table 4: Four-class transcription reliability prediction and cross-domain generalization for MOSS-Audio-8B. We report macro-F1, accuracy, and mean absolute error (MAE) on the in-domain and cross-domain test sets, following the same baselines and evaluation protocol as the Qwen2-Audio-7B-Instruct results in Table 1 . The proposed reliability predictor achieves the strongest overall performance in both settings.
Figure 6: Confusion matrices for transcription reliability prediction across target Audio LLMs. Results are shown for Qwen2-Audio-7B-Instruct, Phi-4-Multimodal, and MOSS-Audio-8B. Rows represent the ground-truth reliability classes, columns represent the predicted reliability classes, and each cell reports the percentage of examples from the corresponding ground-truth class. The accuracy for each target Audio LLM is reported above its confusion matrix.
Parameters
FLOPs
Model
Encoder
Decoder
Head
Encoder
Decoder
Head
Qwen2-Audio-7B-Instruct
636.97M
7.76B
0.33M
1.89T
4.21T
2.12M
Phi-4-Multimodal-Instruct
441.24M
3.84B
0.26M
114.20G
1.32T
1.70M
MOSS-Audio-8B
643.62M
8.19B
329.99K
243.05G
2.67T
2.12M
Appendix
Table 5: Computational overhead of the reliability predictor. We report the parameter counts and FLOPs of the audio encoder, language model decoder, and proposed prediction head. FLOPs are measured with torch.profiler on an approximately 12-second input recording using a single forward pass. Methods relying on generated transcripts or generation uncertainty require executing the decoder, whereas our predictor operates directly on the encoder representations.
Pooling Method
Macro-F1 ↑
Accuracy ↑
MAE ↓
Max Pooling
77.60
79.50
0.22
Mean Pooling
79.69
81.35
0.20
Learned Temporal Pooling
81.10
82.60
0.18
Appendix
Table 6: Ablation of temporal aggregation methods for transcription reliability prediction. All methods operate on frozen Qwen2-Audio-7B-Instruct audio-encoder representations and use the same transcription reliability dataset. Only the temporal aggregation module and classification head are trained.
Transfer setting
Direct Transfer
Aligned Transfer
Alignment method
Global Mean
Global Median
Per-Boundary Mean
Per-Boundary Median
MOSS ← Qwen2
9.69
1.90
1.92
2.30
1.70
MOSS ← Phi-4
12.86
2.71
2.29
2.21
2.05
Mean ↓
11.28
2.31
2.10
2.26
1.88
Appendix
Table 7: Ablation of critical-SNR alignment strategies for cross-model reliability label transfer. We compare direct cross-model transfer with several strategies for aligning source-model critical-SNR boundaries to the target model using 5% of the shared speech–corruption pairs. We report the accuracy drop relative to training with the target model’s own label bank; lower is better.
Figure 7: Transcription reliability information becomes stronger in deeper audio-encoder layers. We probe frozen representations extracted from increasing depths of the Qwen2-Audio-7B-Instruct, Phi-4-Multimodal-Instruct, and MOSS-Audio-8B audio encoders using the same lightweight reliability predictor. Deeper layers generally provide more informative representations, improving macro-F1 and accuracy while reducing MAE.
Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce VoxParadox, an adversarial benchmark with 2,000 verified examples, spanning 10 paralinguistic tasks, created with controlled speech synthesis to intentionally mismatch transcript claims and speaking style, enabling direct measurement of speech paralinguistic understanding. Evaluation of a diverse set of Audio LLMs reveals consistently low accuracy on acoustic ground truth and a strong tendency to follow language-implied (incorrect) answers. To understand the cause of this gap, we perform layer-wise probing and find that (i) paralinguistic cues can degrade in deeper encoder layers and at the encoder--LLM interface, and (ii) even when such cues are available in audio tokens, the language model frequently ignores them. To address these problems, we propose Prompt-Conditioned Layer Mixer (PCLM), which adaptively combines information from multiple audio layers based on the input prompt, and pair it with Direct Preference Optimization (DPO) to explicitly prefer acoustically supported options over language-implied alternatives. These methods substantially improve Audio LLM paralinguistic understanding, improving Audio Flamingo 3 from 17.40% to 65.20% on VoxParadox, and from 37.74% to 54.78% on MMSU paralinguistic subset. Our project page is available at https://voxparadox.github.io/.
Jiacheng Pang, Ashutosh Chaubey, Mohammad Soleymani
Institute for Creative Technologies, University of Southern California, Los Angeles, USA
Large Audio-Language Models show consistent performance gains across speech and audio benchmarks, yet high scores may not reflect true auditory perception. If a model can answer questions without processing the acoustic signal, the benchmark fails as a measure of auditory understanding. We present a diagnostic framework using two axes: text prior, which measures answerability from text and general knowledge alone, and audio reliance, which assesses actual dependency on the acoustic signal. Evaluating eight LALMs across three benchmarks, we find that models retain 60-72% of their full audio scores even without any audio input. Moreover, among items that require audio, only 3.0-4.2% need the complete audio clip; the majority can be resolved using localized fragments. These findings challenge the assumption that benchmark performance equals robust audio understanding, and we conclude with practical guidelines for improving evaluation reliability and benchmark design.
Leonardo Haw-Yang Foo, Chih-Kai Yang, Chen-An Li +2
National Taiwan University · NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
Pooneh Mousavi, Amir Ivry, Mirco Ravanelli +1
Concordia University · Mila – Quebec AI Institute · Technion – IIT +1