Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
Organizations: Beihang University · Didi Chuxing
Abstract
Chain-of-Thought (CoT) empowers Large Language Models (LLMs) to tackle complex problems, but remains constrained by the computational cost and early token commitments in discrete reasoning traces. Recent latent reasoning approaches attempt to optimize efficiency by performing reasoning within continuous hidden states. However, many such methods optimize latent states end to end without a trained interface for intermediate textual readout, and several representative configurations use a pre-defined number of latent steps during inference. In this work, we introduce \textbf{PLaT} (\textbf{P}lanning with \textbf{La}tent \textbf{T}houghts), a framework that decouples latent planning from verbalization. The Planner deterministically evolves latent planning states, while an independent Decoder provides textual readouts when needed. Answer-aware textual stopping allows the latent rollout to use a problem-dependent number of groups rather than a pre-specified chain length. PLaT achieves competitive coverage at larger in several mathematical settings, with lower Pass@1: on Llama-1B GSM8K, it reaches 80.59% Pass@128 versus CODI's 72.37%. These results support PLaT as a candidate-generation interface supplying multiple textual readouts for downstream verification or reranking.
Figures & tables
| Model | Method | GSM8K | GSM-Hard | SVAMP | MultiArith | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| @1 | @32 | @64 | @128 | @1 | @32 | @64 | @128 | @1 | @32 | @64 | @128 | @1 | @32 | @64 | @128 | ||
| GPT-2 | CoT-SFT | 35.30 | 74.20 | 80.41 | 86.33 | 7.10 | 18.35 | 20.65 | 22.52 | 36.63 | 65.37 | 70.50 | 75.47 | 85.00 | 99.07 | 99.44 | 99.63 |
| Coconut | 32.94 | 55.69 | 61.45 | 66.19 | 7.62 | 13.65 | 15.01 | 16.15 | 35.25 | 50.15 | 53.05 | 56.45 | 81.67 | 90.83 | 91.94 | 93.33 | |
| CODI | 40.16 | 63.00 | 67.12 | 69.90 | 8.74 | 14.58 | 15.52 | 16.53 | 38.60 | 53.67 | 55.80 | 58.13 | 92.04 | 96.67 | 96.67 | 96.67 | |
| PLaT | 18.35 | 59.51 | 67.74 | 76.08 | 3.75 | 14.22 | 17.06 | 18.88 | 21.80 | 53.70 | 60.75 | 66.45 | 33.61 | 79.44 | 87.50 | 91.94 | |
| Llama-1B | CoT-SFT | 49.96 | 86.62 | 90.45 | 93.18 | 11.45 | 23.68 | 25.50 | 27.72 | 59.50 | 87.97 | 91.07 | 93.37 | 94.07 | 100.00 | 100.00 | 100.00 |
| CoT | Coconut | CODI | PLaT | |
| Fwd. | 25.55 | 6.00 | 6.00 | 12.83 |
| Time (ms) |
| Pass@1 | Pass@32 | Pass@64 | Pass@128 | |
|---|---|---|---|---|
| PLaT-8 | ||||
| PLaT semantic-step | ||||
| CoT initialization | ||||
| Decoder latent history | ||||
| Decoder question + latent history |
| Method | MATH-500 | AIME 2024 | AIME 2025 | |||
|---|---|---|---|---|---|---|
| Pass@1 | Pass@128 | Pass@1 | Pass@128 | Pass@1 | Pass@128 | |
| CoT-SFT | ||||||
| Coconut | ||||||
| CODI | ||||||
| PLaT-16 | ||||||
| PLaT-32 | ||||||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Definition | Note |
|---|---|---|
| Input question sequence | Token sequence | |
| Complete reasoning chain | segments, tokens | |
| The -th textual segment | At most tokens | |
| Index of textual segments | ||
| Index within a latent group | ||
| Planner backbone | LoRA-adapted |
| Method | Dataset | Pass@1 | Pass@32 | Pass@64 | Pass@128 |
|---|---|---|---|---|---|
| CoT-SFT | GSM8K | ||||
| CoT-SFT | GSM-Hard | ||||
| CoT-SFT | SVAMP | ||||
| CoT-SFT | MultiArith | ||||
| Coconut | GSM8K | ||||
| Coconut | GSM-Hard |
| Method | Dataset | Pass@1 | Pass@32 | Pass@64 | Pass@128 |
|---|---|---|---|---|---|
| CoT-SFT | GSM8K | ||||
| CoT-SFT | GSM-Hard | ||||
| CoT-SFT | SVAMP | ||||
| CoT-SFT | MultiArith | ||||
| Coconut | GSM8K | ||||
| Coconut | GSM-Hard |
| Method | Dataset | Pass@1 | Pass@32 | Pass@64 | Pass@128 |
|---|---|---|---|---|---|
| CoT-SFT | GSM8K | ||||
| CoT-SFT | GSM-Hard | ||||
| CoT-SFT | SVAMP | ||||
| CoT-SFT | MultiArith | ||||
| Coconut | GSM8K | ||||
| Coconut | GSM-Hard |
| Pass@1 | Pass@32 | Pass@64 | Pass@128 | |
|---|---|---|---|---|
| CoT initialization | ||||
| semantic-step | ||||
| PLaT-2 | ||||
| PLaT-5 | ||||
| PLaT-8 | ||||
| PLaT-8, |
| Pass@1 | Pass@32 | Pass@64 | Pass@128 | |
|---|---|---|---|---|
| CoT initialization | ||||
| semantic-step | ||||
| PLaT-2 | ||||
| PLaT-5 | ||||
| PLaT-8 | ||||
| PLaT-8, |
| Pass@1 | Pass@32 | Pass@64 | Pass@128 | |
|---|---|---|---|---|
| CoT initialization | ||||
| semantic-step | ||||
| PLaT-2 | ||||
| PLaT-5 | ||||
| PLaT-8 | ||||
| PLaT-8, |
| Pass@1 | Pass@32 | Pass@64 | Pass@128 | |
|---|---|---|---|---|
| CoT initialization | ||||
| semantic-step | ||||
| PLaT-2 | ||||
| PLaT-5 | ||||
| PLaT-8 | ||||
| PLaT-8, |
| Human label | Valid | Invalid | Ambiguous |
|---|---|---|---|
| Valid | 170 | 9 | 2 |
| Invalid | 5 | 14 | 0 |
| Metric | Cohen’s Kappa ( ) | Accuracy |
|---|---|---|
| Value | 0.5943 | 0.9200 |