Audio-language models (ALMs) can exploit textual shortcuts to answer questions while overlooking acoustic evidence, weakening audio understanding. On-policy distillation (OPD) trains compact ALMs by supervising student-generated responses with teacher predictions, but does not explicitly distinguish acoustic support from linguistic predictability. We propose Reward-Tilted On-Policy Distillation (RT-OPD) to strengthen acoustic grounding. Given the same question and student-generated text, a frozen teacher predicts the next token with and without audio inputs. Their log-probability contrast defines a reward that reshapes the teacher distribution for reverse-KL distillation, emphasizing the additional evidence provided by audio. Across two compact students and three benchmarks, RT-OPD consistently outperforms Vanilla OPD. Experiments with silenced and replacement audio further suggest that RT-OPD strengthens the student's reliance on acoustic evidence. Our 3B model achieves 72.72% accuracy on MMAU, the highest among the compared 3B models and competitive with several 7B and 8B models. Code and model checkpoints are available at https://github.com/KaiyangLi1992/RT-OPD.
Figures & tables
Model
Reported size
MMAU
Audio-Thinker [ 9 ]
8.4B
75.98
Nova 2 Omni [ 10 ]
—
75.28
Step-Audio-2 [ 11 ]
—
73.86
Mizar-3B (ours)
3B
72.72
MiMo-Audio [ 12 ]
7B
72.59
Audio Flamingo 3 [ 13 ]
8.2B
72.42
Table 1: ALM performance on MMAU (%), ranked by accuracy. External scores and reported sizes follow the official leaderboard [ 8 ] . Mizar-3B 1 1 1 The name “Mizar” was inspired by the star used for navigation. We hope our model and method can provide guidance for future research and engineering optimization on audio-language models. is obtained by post-training Ke-Omni-R-3B with RT-OPD; its score is averaged over five seeds. Protocols and parameter-count conventions differ across models.
Method
MMAU
MMAR
ADQA-cl.
Macro-3
Ke-Omni-R-7B (teacher)
Teacher
73.86
61.70
54.85
63.47
Ke-Omni-R-3B
Student
68.70
52.70
49.59
57.00
CE only
71.32
58.60
54.41
61.44 ± 0.34
Vanilla OPD [ 4 ]
72.21
57.78
56.02
62.00 ± 0.31
Table 2: Main results (%). Trained methods: five-seed means, with Macro-3 as mean ± sample SD over seeds; teacher and students before distillation: single runs. Bold: best mean per student block.
Variant
MMAU
MMAR
ADQA-cl.
Macro-3
Ke-Omni-R-3B
RT-OPD
72.72
60.08
56.45
63.08 ± 0.36
Unmatched
72.84
59.06
56.09
62.66 ± 0.51
Sharpening
72.59
56.95
54.57
61.37 ± 0.40
Model ref.
72.57
58.10
53.44
61.37 ± 0.15
Off-policy
71.72
58.84
56.17
62.24 ± 0.30
Table 3: Target and rollout ablations (%). Macro-3 shows mean ± sample SD over per-seed scores. Bold marks the best mean within each student block.
Method
Orig.
Silence
Replaced
MMAU
Vanilla OPD
72.21
52.20 ( − 20.01)
48.56 ( − 23.65)
RT-OPD
72.72
52.03 ( − 20.69)
47.07 ( − 25.65)
MMAR
Vanilla OPD
57.78
38.62 ( − 19.16)
38.74 ( − 19.04)
RT-OPD
60.08
37.62 ( − 22.46)
38.44 ( − 21.64)
Table 4: Audio perturbation on Ke-based 3B models (five-seed mean accuracy, %). Parentheses show changes from the original audio (percentage points).