cs.SDSep 23, 2026

Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models

Authors: Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji

Organizations: NEC Laboratories America, Inc., USA · School of Computing, University of Connecticut, USA

Abstract

Audio-language models (ALMs) can exploit textual shortcuts to answer questions while overlooking acoustic evidence, weakening audio understanding. On-policy distillation (OPD) trains compact ALMs by supervising student-generated responses with teacher predictions, but does not explicitly distinguish acoustic support from linguistic predictability. We propose Reward-Tilted On-Policy Distillation (RT-OPD) to strengthen acoustic grounding. Given the same question and student-generated text, a frozen teacher predicts the next token with and without audio inputs. Their log-probability contrast defines a reward that reshapes the teacher distribution for reverse-KL distillation, emphasizing the additional evidence provided by audio. Across two compact students and three benchmarks, RT-OPD consistently outperforms Vanilla OPD. Experiments with silenced and replacement audio further suggest that RT-OPD strengthens the student's reliance on acoustic evidence. Our 3B model achieves 72.72% accuracy on MMAU, the highest among the compared 3B models and competitive with several 7B and 8B models. Code and model checkpoints are available at https://github.com/KaiyangLi1992/RT-OPD.

Figures & tables

Explore similar work

CardsList
  1. AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

    Sep 30, 2026Pooneh Mousavi, Amir Ivry, Mirco Ravanelli +1Large Audio Language ModelsHallucination Mitigation

  2. Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

    Sep 23, 2026Kaiyang Li, Shaobo Han, Yue Tian +1Large Audio Language ModelsOpencode