cs.SDOct 3, 2026

AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models

Authors: Lee Seung-woo, Bowen Qi

Abstract

Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same computational depth on both. We introduce AdaLoop, a lightweight recurrent module that learns how many latent refinement steps a given audio--question pair requires. A shared transformer block iterates over the audio representation, guided by the question, while a learned halting mechanism exits the loop once the representation is ready. AdaLoop adds fewer than 3% of the base model's parameters and plugs into any audio encoder--language model pair without modifying either component. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, AdaLoop raises the average accuracy by 2.9 to 3.8 points, with the largest gains on perception-heavy subtasks where the model learns to apply deeper reasoning.

Explore similar work

CardsList
  1. SkillFormer: Skill-Decomposed Adaptation for Audio Language Models

    Oct 6, 2026Lee Seung-woo, Bowen Qi, Kim Min-jun +1Multi-Task LearningAudio Understanding

  2. Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

    Sep 23, 2026Kaiyang Li, Shaobo Han, Yue Tian +1Audio QAAudio-Language Models