Audio-language models (ALMs) can exploit textual shortcuts to answer questions while overlooking acoustic evidence, weakening audio understanding. On-policy distillation (OPD) trains compact ALMs by supervising student-generated responses with teacher predictions, but does not explicitly distinguish acoustic support from linguistic predictability. We propose Reward-Tilted On-Policy Distillation (RT-OPD) to strengthen acoustic grounding. Given the same question and student-generated text, a frozen teacher predicts the next token with and without audio inputs. Their log-probability contrast defines a reward that reshapes the teacher distribution for reverse-KL distillation, emphasizing the additional evidence provided by audio. Across two compact students and three benchmarks, RT-OPD consistently outperforms Vanilla OPD. Experiments with silenced and replacement audio further suggest that RT-OPD strengthens the student's reliance on acoustic evidence. Our 3B model achieves 72.72% accuracy on MMAU, the highest among the compared 3B models and competitive with several 7B and 8B models. Code and model checkpoints are available at https://github.com/KaiyangLi1992/RT-OPD.
Figures & tables
Model
Reported size
MMAU
Audio-Thinker [ 9 ]
8.4B
75.98
Nova 2 Omni [ 10 ]
—
75.28
Step-Audio-2 [ 11 ]
—
73.86
Mizar-3B (ours)
3B
72.72
MiMo-Audio [ 12 ]
7B
72.59
Audio Flamingo 3 [ 13 ]
8.2B
72.42
Table 1: ALM performance on MMAU (%), ranked by accuracy. External scores and reported sizes follow the official leaderboard [ 8 ] . Mizar-3B 1 1 1 The name “Mizar” was inspired by the star used for navigation. We hope our model and method can provide guidance for future research and engineering optimization on audio-language models. is obtained by post-training Ke-Omni-R-3B with RT-OPD; its score is averaged over five seeds. Protocols and parameter-count conventions differ across models.
Method
MMAU
MMAR
ADQA-cl.
Macro-3
Ke-Omni-R-7B (teacher)
Teacher
73.86
61.70
54.85
63.47
Ke-Omni-R-3B
Student
68.70
52.70
49.59
57.00
CE only
71.32
58.60
54.41
61.44 ± 0.34
Vanilla OPD [ 4 ]
72.21
57.78
56.02
62.00 ± 0.31
Table 2: Main results (%). Trained methods: five-seed means, with Macro-3 as mean ± sample SD over seeds; teacher and students before distillation: single runs. Bold: best mean per student block.
Variant
MMAU
MMAR
ADQA-cl.
Macro-3
Ke-Omni-R-3B
RT-OPD
72.72
60.08
56.45
63.08 ± 0.36
Unmatched
72.84
59.06
56.09
62.66 ± 0.51
Sharpening
72.59
56.95
54.57
61.37 ± 0.40
Model ref.
72.57
58.10
53.44
61.37 ± 0.15
Off-policy
71.72
58.84
56.17
62.24 ± 0.30
Table 3: Target and rollout ablations (%). Macro-3 shows mean ± sample SD over per-seed scores. Bold marks the best mean within each student block.
Method
Orig.
Silence
Replaced
MMAU
Vanilla OPD
72.21
52.20 ( − 20.01)
48.56 ( − 23.65)
RT-OPD
72.72
52.03 ( − 20.69)
47.07 ( − 25.65)
MMAR
Vanilla OPD
57.78
38.62 ( − 19.16)
38.74 ( − 19.04)
RT-OPD
60.08
37.62 ( − 22.46)
38.44 ( − 21.64)
Table 4: Audio perturbation on Ke-based 3B models (five-seed mean accuracy, %). Parentheses show changes from the original audio (percentage points).
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
Pooneh Mousavi, Amir Ivry, Mirco Ravanelli +1
Concordia University · Mila – Quebec AI Institute · Technion – IIT +1
While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X3-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditioned on its own acoustic perception, while the teacher provides token-level guidance using matched textual inputs and verified answers. We further construct a three-tier symmetric corpus covering textual reasoning rendered into speech, audio-event reasoning grounded in complex acoustic scenes, and spoken-dialogue reasoning involving paralinguistic cues. This design extends cross-modal distillation beyond textually recoverable content to reasoning grounded in non-linguistic events, prosody, and conversational context. Experiments on MMSU, MMAU, BIG Bench Audio, and MMAR demonstrate that X3-OPD substantially improves audio-grounded reasoning and chain-of-thought quality while largely preserving the model's existing capabilities under domain shift.
Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M parameters. We introduce a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM. Its architecture connects a compact CED-Small audio encoder to SmolLM2-135M through a frequency-merging mapper. With supervision drawn from ReasonAQA, AudioMCQ, and AVQA, the model undergoes three training stages: audio-language alignment (Stage 1), audio-dependent fine-tuning (Stage 2), and post-training (Stage 3) aimed at strengthening weak skills while retaining learned capabilities. Across five random seeds, Mizar achieves mean accuracies of 52.92% on MMAU, 42.42% on MMAR, and 36.02% on ADQA-clean, surpassing the previous best-performing ALM below 200M parameters on all three benchmarks. It also supports local inference on a single CPU: on questions from the MMAU benchmark, the mean latency from opening the audio file to generating a complete answer is 1.09 seconds. Code and checkpoints are available at https://github.com/KaiyangLi1992/Mizar_159M.
Kaiyang Li, Shaobo Han, Yue Tian +1
NEC Laboratories America, Inc, USA · School of Computing, University of Connecticut, USA