cs.CLOct 1, 2026

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

Authors: Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, +7 more

Organizations: The Chinese University of Hong Kong, Shenzhen · Tencent Hunyuan · Tsinghua University · The Hong Kong University of Science and Technology · Amphion Technology Co., Ltd.

Abstract

Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

    Sep 29, 2026Shengbo Cai, Yuxiang Wang, Jingran Xie +5Speech Language ModelsParalinguistic Cues

  2. Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models

    Apr 20, 2026Xiang He, Chenxing Li, Jinting Wang +5Audio UnderstandingText-To-Audio

  3. Learning When to Think While Listening in Large Audio-Language Models

    May 26, 2026Zhiyuan Song, Weici Zhao, Yang Xiao +3Large Audio Language Models