AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
Organizations: The Chinese University of Hong Kong, Shenzhen · Tencent Hunyuan · Tsinghua University · The Hong Kong University of Science and Technology · Amphion Technology Co., Ltd.
Abstract
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.
Figures & tables
| EchoMind | Emotion | General Ability | VoiceBench | ||||||||
| Method | MCQ | OpenQ | IEMOCAP | MELD | MMSU | MMAU-Pro | GPQA | OBQA | MMSU | BBH | IFEval |
| Qwen2.5-Omni based methods | |||||||||||
| Base | 58.09 | 3.492 | 66.24 | 52.87 | 61.60 | 55.66 | 23.44 | 79.78 | 51.53 | 66.70 | 54.64 |
| CoT | 67.94 | 3.631 | 71.39 | 59.12 | 63.72 | 62.33 | 35.35 | 85.27 | 61.97 | 68.30 | 55.30 |
| 60.83 | 3.548 | 67.12 | 54.29 | 62.08 | 57.41 | 25.82 | 80.66 | 53.87 | 64.50 | 53.47 | |
| 62.37 | 3.532 | 67.85 | 54.06 | 62.40 | 57.82 | 28.21 | 81.32 | 55.30 | 64.20 | 53.92 | |
| Variant | EchoMind MCQ | MMSU Acc. | Latent states | EOL hit rate |
|---|---|---|---|---|
| AURAL-RL | 68.63 | 64.76 | 12.36 | 100.00 |
| w/o CSA | 64.74 | 63.16 | 12.58 | 96.62 |
| w/o LSS | 66.82 | 64.30 | 11.92 | 98.80 |
| w/o conciseness bonus | 68.81 | 63.96 | 14.77 | 100.00 |
| 65.99 | 63.68 | 13.15 | 100.00 | |
| 67.02 | 63.84 | 11.94 | 100.00 |
| Generation scheme | Two valid paths | Individual acc. | GPQA | OBQA | BBH |
|---|---|---|---|---|---|
| Standard AURAL-RL, one path | – | 63.57% | 39.38% | 87.25% | 66.00% |
| One path per GMM component | 38.88% | 62.19% | 42.49% | 89.23% | 68.30% |
| Four samples from one component | 12.84% | 62.81% | 40.29% | 87.69% | 66.50% |
| Four components merged | 16.09% | 62.48% | 40.84% | 88.13% | 67.00% |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | CoT initialization | AURAL-SFT | AURAL-RL | CoT-RL |
|---|---|---|---|---|
| Epochs | 1 | 1 | 2 | 2 |
| Updates | 10,700 | 10,700 | 484 | 484 |
| Nodes | 2 | 2 | 2 | 2 |
| GPUs | 16 | 16 | 16 | 16 |
| Batch per GPU | 4 | 4 | 1 | 1 |
| Gradient accumulation | 1 | 1 | 1 | 1 |
| Setting | Value |
|---|---|
| Latent structure | Compression , chunk size , components |
| Head width | 3,584 |
| Covariance | Low-rank plus diagonal, rank |
| Log standard deviation | Clipped to |
| Loss weights | Token CE 1, chunk NLL 1, second-pass LSS 1 |
| CSA | Soft targets over the tokens in each pooled CoT group |
| Setting | Value |
|---|---|
| Clipping coefficient | 0.2 |
| Advantage constant | |
| Action loss weights | Latent chunk 1, EOL decision 1, answer token 1 |
| Sampling temperatures | Latent 1.0, EOL 1.0, answer 1.0 |
| Answer sampling | Top , maximum 128 tokens |
| Latent budget | Maximum 64 states |
| Benchmark | Subset | Examples | Metric |
|---|---|---|---|
| EchoMind | MCQ | 13,401 | Accuracy |
| EchoMind | OpenQ | 4,715 | Mean judge score |
| MMSU | Full benchmark | 5,000 | Accuracy |
| MMAU-Pro | Objectively scored subset | 4,680 | Accuracy |
| GPQA | Main, Diamond, and Extended | 546 | Accuracy |
| VoiceBench | OpenBookQA | 455 | Accuracy |
| System | Decoding |
|---|---|
| AURAL | Latent and EOL temperatures 1.0, maximum 100 latent states, greedy visible answer |
| Qwen base | Temperature 0, top |
| Kimi base | Text and audio temperatures 0, top , repetition penalty 1.0, repetition window 16 |
| System | Mean | p50 | p95 | Speedup |
| Time to first visible answer token | ||||
| Base | 0.050 | 0.049 | 0.069 | |
| CoT-RL | 1.218 | 1.090 | 2.056 | Reference |
| AURAL-SFT ( ) | 0.107 | 0.106 | 0.131 | |
| AURAL-RL ( ) | 0.103 | 0.103 | 0.121 | |
| AURAL-RL ( ) | 0.087 | 0.086 | 0.105 | |
| Operator | Answer change | Branch accuracy | DCR | ||||
|---|---|---|---|---|---|---|---|
| Mode intervention | 0.312 | 0.283 | 0.518 | 3.41 | 38.68 | 62.19 | 38.88 |
| Within-component | 0.087 | 0.064 | 0.791 | 1.28 | 14.24 | 62.81 | 12.84 |
| Collapsed mixture | 0.108 | 0.089 | 0.742 | 1.54 | 18.49 | 62.48 | 16.09 |
| Single Gaussian | 0.071 | 0.047 | 0.816 | 0.93 | 11.34 | 61.77 | 9.45 |
| GPQA | OpenBookQA | BBH | ||||
| Operator | Branch | Vote | Branch | Vote | Branch | Vote |
| AURAL-RL, single rollout | 39.38 | n/a | 87.25 | n/a | 66.00 | n/a |
| Mode intervention | 38.19 | 42.49 | 85.82 | 89.23 | 64.55 | 68.30 |
| Within-component | 38.83 | 40.29 | 86.37 | 87.69 | 65.18 | 66.50 |
| Collapsed mixture | 38.46 | 40.84 | 86.04 | 88.13 | 64.88 | 67.00 |
| Single Gaussian | 36.72 | 38.28 | 84.89 | 86.37 | 64.93 | 65.90 |
| Entropy tercile | range | Answer change | DCR | ||
|---|---|---|---|---|---|
| Bottom | 0.189 | 0.158 | 21.44 | 22.79 | |
| Middle | 0.314 | 0.284 | 38.53 | 39.13 | |
| Top | 0.434 | 0.407 | 56.07 | 54.72 |
| Branch | Chinese | English | Total | Hours |
|---|---|---|---|---|
| LIME-440K | 197K | 123K | 320K | 367 |
| EmotionCoT-35K | 0 | 18K | 18K | 22 |
| HumanSpeech-1M | 47K | 28K | 75K | 83 |
| GeneralSpeech | 133K | 137K | 270K | 526 |
| Total | 377K | 306K | 683K | 1,000 |
| Source | Retained | Acceptance |
|---|---|---|
| LIME-440K | ||
| LIME Core Chinese | 191,405 | 89.9% |
| LIME Core English | 89,785 | 93.5% |
| ECD-TSE extension | 24,991 | 29.8% |
| Emotion Speech Dataset extension | 15,169 | 59.4% |
| Total | 321,350 | 76.8% |
| Language | Source | Clips | Hours |
|---|---|---|---|
| English | Affective media collection | 259,152 | 337.10 |
| English | GigaSpeech Audiobook | 254,871 | 303.79 |
| English | GigaSpeech YT | 130,254 | 187.28 |
| English | GigaSpeech POD | 108,657 | 156.90 |
| English | VoxMovies | 3,752 | 3.43 |
| Chinese | WenetSpeech L | 511,155 | 542.48 |
| Language | Source | Retained |
|---|---|---|
| Chinese | WenetSpeech L | 26,854 |
| Chinese | Affective media collection | 19,784 |
| English | GigaSpeech Audiobook | 9,464 |
| English | GigaSpeech YT | 6,809 |
| English | Affective media collection | 6,308 |
| English | GigaSpeech POD | 4,681 |
| Task family | Share |
|---|---|
| Explanation and comprehension | 35.8% |
| Numeric and mathematical | 20.2% |
| Limited generation | 20.0% |
| Multiple choice | 13.6% |
| Factual entity | 8.1% |
| Boolean judgement | 2.3% |
| State | Content |
|---|---|
| Dialogue | |
| Cue | Audible properties that affect the reply, such as a shaky voice and long pauses |
| Need | The conversational need suggested by those cues, such as reassurance before advice |
| Risk | Ambiguity or safety considerations that affect the reply |
| Plan | The response act, order, and tone, such as brief validation followed by one next step |
| Emotion classification | |
| Branch | p10 | p50 | p90 | Mean states |
|---|---|---|---|---|
| LIME-440K | 50 | 71 | 95 | 3.67 |
| EmotionCoT-35K | 41 | 46 | 52 | 3.00 |
| HumanSpeech-1M | 49 | 65 | 89 | 3.31 |
| GeneralSpeech | 40 | 63 | 88 | 2.10 |
| Pairing | Fully executed | Partly executed | Contradictory |
|---|---|---|---|
| True pair | 91.3% | 8.0% | 0.7% |
| Mismatched control | 6.0% | 31.0% | 63.0% |
| Constructed CoT | Response |
|---|---|
| D1. Good news, a heavy voice LIME-440K | |
| Transcript: “They promoted me to lead the sustainability project…” | |
| Cue: low pitch, slow tempo, heavy sighing breath at the start, downward intonation heavy reluctance, overwhelm Need: validation that a promotion can feel like a burden rather than a reward Risk: congratulating them forces a positive reaction they do not feel Plan: acknowledge the heavy weight of the news, then ask what makes this feel daunting | That sounds like a massive undertaking, and from your tone, it feels more like a heavy weight than a celebration. What is the part of this new role that is weighing on you the most right now? |
| D2. Letting a success be a success LIME-440K | |
| Transcript: “The professor approved my thesis draft without any changes!” | |
| Cue: bright, rapid tempo, rising pitch, audible smile relief, pride Need: validation of massive effort behind clean pass Risk: focusing on next steps ruins rare moment of absolute victory Plan: match triumphant energy, highlight zero-change approval, invite celebration | Without any changes? That is practically unheard of! You must be walking on air right now after all that work. How are you going to celebrate tonight? |
| Constructed CoT | Response |
|---|---|
| S1. Amusement in a factual utterance Expresso | |
| Transcript: “There’s an alarm set for today at two thirty AM, another for tomorrow at two fifteen PM, and for Wednesday at seven PM.” | |
| Cue: breathy giggling, wide pitch variation, and bright, smiling vocal resonance. Rule out: confused — presence of bubbling laughter and rising, playful pitch contours. State: high-arousal positive valence with spontaneous amusement. | happy |
| S2. Strained restraint versus overt anger IEMOCAP | |
| Transcript: “Okay, no, look, I- I- I- I’m really – I’m trying to be patient. I know this isn’t your fault. I know that you didn’t do it” | |
| Cue: rapid and uneven tempo, audible sighs, tense vocal quality, frequent self-interruptions Rule out: angry — lacks the high-energy acoustic projection and sharp volume spikes State: moderate-to-high arousal negative valence, showing signs of strained self-control | frustrated |
| G1. Minimum feasible dividend GeneralSpeech | |
| Question: When two numbers are divided, the quotient is 12 and the remainder is 3. What is the smallest possible dividend? | |
| Original CoT: Qwen3-Omni-Thinking | Our rewritten CoT |
| Okay, let’s see. The problem says when two numbers are divided, the quotient is 12 and the remainder is 3. We need to find the smallest possible dividend. Hmm, right. So first, I remember that when you divide numbers, there’s a relationship between the dividend, divisor, quotient, and remainder. Let me recall the formula: Dividend = Divisor Quotient + Remainder. Yeah, that’s right. So if we let the divisor be , the quotient is 12, remainder is 3, then the dividend is . But wait, the remainder has to be less than the divisor, right? Because if the remainder was equal to or larger than the divisor, you could divide again. So the remainder must satisfy . In this case, the remainder is 3, so . That means the divisor has to be greater than 3. So the smallest possible divisor would be 4, since it has to be an integer greater than 3. Let me check that. If , then the dividend would be . Let’s verify: 51 divided by 4 is 12 with a remainder of 3, because 4 times 12 is 48, and 51 minus 48 is 3. Yep, that works. Is there a smaller divisor? Well, if , then the remainder would have to be less than 3, but the remainder here is 3, which is not less than 3, so that’s not allowed. If , same problem, remainder 3 is bigger than 2, so you could divide more. So the smallest possible divisor is 4, which gives the smallest dividend of 51. Let me just make sure. Let’s try : , but in integer division, it’s 12 with remainder 3. Correct. If we take , then dividend is , which is bigger, so 51 is smaller. So yeah, 51 is the smallest possible dividend. | Given: Dividend = Divisor , and remainder 3 must be less than divisor. Step: The smallest integer divisor satisfying is 4. Step: Dividend . Response The smallest possible dividend is 51, because the divisor must be at least 4, and 12 times 4 plus 3 equals 51. |
| G2. Selecting the obtuse angle GeneralSpeech | |
| Question: The ratio of two interior angles of a parallelogram is . What is the degree measure of one of the two obtuse angles of the parallelogram? | |
| Original CoT: Qwen3-Omni-Thinking | Our rewritten CoT |
| Okay, let’s try to figure out this problem. So, we have a parallelogram, and the ratio of two interior angles is . We need to find the measure of one of the obtuse angles. Hmm, first, I remember that in a parallelogram, opposite angles are equal, and consecutive angles are supplementary. That means if one angle is, say, , the next one is , right? Because they add up to 180 degrees. So, the problem says the ratio of two interior angles is . Let’s assume these are consecutive angles because if they were opposite angles, they would be equal, so their ratio would be , which isn’t . So, they must be consecutive angles. Let me confirm that. In a parallelogram, adjacent angles are supplementary, so they add up to 180 degrees. So, if the ratio is , let’s let the measures be and . Then, degrees. Because they are consecutive angles, they should add up to 180. Let me check that. Yes, in a parallelogram, consecutive angles are supplementary. So, . That would be , so . Then, the angles would be degrees and degrees. Now, which one is obtuse? Obtuse angles are greater than 90 degrees, so 100 degrees is obtuse. So, the measure of one of the obtuse angles is 100 degrees. Let me just make sure. In a parallelogram, opposite angles are equal. So, if two consecutive angles are 80 and 100, then the other two angles are also 80 and 100. So, there are two acute angles (80 degrees) and two obtuse angles (100 degrees). So, the obtuse angles are 100 degrees each. That makes sense. Let me verify the ratio: simplifies to , which matches the problem statement. Yep, that works. So, the answer should be 100 degrees. | Given: parallelogram interior angles ratio , opposite angles equal, consecutive supplementary. Step: The ratio must be for consecutive angles. Let them be and . Then . Step: The two angles are and . The larger is obtuse. Response The obtuse angle is 100 degrees. |
| G3. Resolving two stars in blue light GeneralSpeech | |
| Question: Why might some stars appear double in blue light, though they could not be resolved in red light with the same telescope? | |
| Original CoT: Qwen3-Omni-Thinking | Our rewritten CoT |
| We are considering the phenomenon where a star appears double in blue light but not in red light with the same telescope. The key point is the diffraction limit of the telescope. The angular resolution ( ) of a telescope is given by the Rayleigh criterion: where: = wavelength of light = diameter of the telescope’s aperture Since blue light has a shorter wavelength ( ) than red light ( ), the angular resolution for blue light is better (smaller ) than for red light ( ). Therefore, if two stars are separated by an angular distance that is greater than (so they are not resolved in red light) but less than (so they are resolved in blue light), then we would see them as double in blue light but not in red. However, note that the Rayleigh criterion is an approximation and the actual resolution also depends on the point spread function and other factors. But the main idea is that the resolving power of a telescope improves as the wavelength decreases. Example: Let nm, nm. Then . So the resolving power in blue is about 1.75 times better than in red. Therefore, if two stars are separated by an angle that is between and , they will be resolved in blue but not in red. Hence, the answer is: because blue light has a shorter wavelength, which allows the telescope to resolve finer details (smaller angular separations) than red light. So the phenomenon occurs due to the wavelength dependence of the diffraction limit. | Given: telescope angular resolution Step: blue light has shorter than red light smaller for blue Step: if two stars have angular separation such that , they are unresolved in red but resolved in blue Response Stars appear double in blue light because the telescope’s diffraction limit depends on wavelength. Blue light, with its shorter wavelength, gives better angular resolution than red light. If the stars’ separation falls between the red and blue resolution limits, they blend into one in red but appear as two in blue. |