cs.CLSep 29, 2026

Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

Authors: Shengbo Cai, Yuxiang Wang, Jingran Xie, Zhisheng Zhang, Shun Lei, Di Cao, Teddy Sun, Zhiyong Wu

Organizations: Tsinghua University · Tencent Hunyuan · The Chinese University of Hong Kong, Shenzhen

Abstract

Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.

Figures & tables

Explore similar work

CardsList
  1. AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

    Oct 1, 2026Yuxiang Wang, Kunyu Feng, Yuancheng Wang +12Speech Language ModelsEfficient Latent Reasoning

  2. ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models

    Jun 9, 2026Yuxiang Wang, Qinke Ni, Shengbo Cai +3Paralinguistic CuesSpeech Language Models

  3. RetroThinker: Enabling Retrospective Thinking in Speech LLMs

    Sep 12, 2026Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed +1LLM Reasoning StrategiesSpeech Language Models