Visual latent reasoning compresses rendered derivations into compact intermediate states, reducing textual reasoning overhead. Existing approaches differ in how they represent these states: continuous methods avoid vocabulary constraints, whereas discrete methods improve accuracy through quantization into a finite codebook. Our analysis of representative continuous and discrete systems identifies two functional requirements: answers must rely on latent states, and those states must carry valid, problem-specific reasoning. Continuous latents influence answers despite collapsed reasoning content, whereas discrete latents retain recoverable intermediate reasoning that answer prediction largely bypasses. To address these challenges, we propose Continuous Anchored Latent Reasoning (CALR), which connects latent formation with answer use through functional anchoring. With reference latents from information-balanced compression, CALR couples latent-mediated answer supervision with derivation-level semantic anchoring: the former routes answer supervision through intermediate states, while the latter grounds their decoded content in problem-specific derivations. A parallel-to-autoregressive curriculum develops sequential reasoning by conditioning subsequent latent blocks on generated prefixes. Evaluations on five mathematical reasoning benchmarks across model families show substantial accuracy gains. Under matched budgets, CALR gains 26.0 percentage points over a comparable continuous latent reasoning method. Further analyses show that its latents support answer prediction and carry problem-specific intermediate reasoning.
Figures & tables
Figure 1: The latent handoff challenge. (a) Continuous latents can support answer prediction without carrying problem-specific reasoning. (b) Discrete latents can encode recoverable reasoning yet be bypassed by the answer model. (c) CALR couples latent reliance (C1) with derivation anchoring for reasoning capability (C2).
Figure 2: Complementary handoff failures in continuous and discrete latents. (a) Both favor the question, with higher latent attention in the continuous system. (b) Content replacement has little effect ( Λcontent≈0 ); interface disruption yields accuracy drops of 0.42 and 0.015 , respectively.
Figure 3: CALR training framework: functional anchoring followed by parallel drafting and autoregressive refinement. Purple dashed paths denote the unified internal-reader route.
Resampling
GSM8K-Aug
Sim. (%) ↑
CE ↓
Global cross-attn.
92.8
0.183
Fixed anchors
95.3
0.158
Dynamic anchors
99.1
0.107
Table 5: Resampler ablations.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Parallel
Autoregressive
Reader question mask pr
1.0/0.5/0.3 by epoch
0.3
Gradient gate α
0→1 in epoch 2
0→1 over 3,000 steps
External weight w
1→0.5 in epoch 2
1→0.5 over 3,000 steps
Readiness checks
–
Every 500 steps; 200 problems
Reopening window
–
Steps 6,000–8,000
Appendix
Table 6: Unified-reader schedules. Epochs and steps are local to each phase.
Source
Training use
Examples
GSM8K-Aug
Resampler, reader, and producer
384,618
MATH
Resampler
11K
MathX-5M subset
Resampler
818K
Appendix
Table 7: Training data and their roles in CALR.
Hyperparameter
Stage I: Representation
Stage I: Readout
Stage II: Drafting
Stage II: Refinement
Training data
Rendered chains (Appendix B.2 )
GSM8K-Aug 385K
GSM8K-Aug 385K
GSM8K-Aug 385K
Optimizer / LR schedule
AdamW / cosine
AdamW / cosine
AdamW / cosine
AdamW / cosine
Peak learning rate
10−4
10−4
10−4
10−4
Warm-up steps
300
200
360
360
Steps (epochs)
12,000 ( ≈1 )
6,000 (2)
18,000 (3)
18,000 (3)
Global batch size
128
128
64
64
Appendix
Table 8: Per-stage training configurations. Rows marked “unified” apply only to unified CALR; the decoupled form uses the same shared settings with α=0 , w=1 , and no internal reader.
System
Question attention
Latent attention
Latent/question gradient ratio
RoT
0.51
0.33
1.12
DLR-SFT
0.72
0.05
0.13
Appendix
Table 9: Answer-stage attention and gradient measurements.
Condition
Metric
Set 1
Set 2
Feedback substitution + suffix regeneration
Expected
75.0
80.0
Reader-input substitution only
Expected
0.8
0.8
Same intermediate, different answer
Orig.
87.5
82.5
Appendix
Table 17: Intermediate-state substitution on two disjoint sets of 120 problems each. Expected denotes the percentage of outputs matching yexp ; Orig. denotes accuracy against the original answer y . Regeneration retains the original first two blocks at readout. All values are percentages.
Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and extended inference time. Latent reasoning reduces this cost by shifting computation into a latent space; however, continuous latent methods are hard to train, suffering from unstable and uninterpretable reasoning trajectories. We argue these issues stem from a misalignment between continuous-space reasoning and discrete symbolic supervision, as continuous states lack explicit anchors for step-by-step alignment. To resolve this, we propose \textbf{Discrete Latent Reasoning~(DLR)}, the first method that converts continuous latent states into explicit discrete tokens. Inspired by render-based compression, we render textual chains of thought into images, extract visual features, and construct a discrete latent vocabulary via clustering-based fine-tuning. Expanding the vocabulary and output head enables standard autoregressive modeling over both natural language and latent tokens, supporting pretraining alignment, SFT, and RL. Experiments on five reasoning benchmarks and two model series~(Qwen3-VL and LLaMA-3) confirm that \textbf{DLR} outperforms prior latent reasoning baselines with up to \textbf{20× compression}. Furthermore, the learned latent trajectories retain an interpretable semantic structure. Overall, discrete latent tokens provide a controllable and interpretable basis for efficient latent reasoning.
Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost due to lengthy autoregressive traces. Existing latent reasoning methods offer a promising alternative, yet they often treat reasoning as uniformly compressible, causing precision-critical intermediate steps to be overly compressed and thereby degrading reasoning accuracy. In this work, we propose Selective Latent Thinking (SLT), a framework that selectively compresses redundant reasoning spans into latent representations while preserving precision-critical spans as explicit CoT within the same reasoning trajectory. Specifically, SLT first uses a lightweight decoder to anticipate a short upcoming reasoning span, and then applies confidence-based gating to determine the longest span that can be reliably compressed. The accepted span is encoded into a compact latent representation to improve reasoning efficiency, while uncertain or precision-critical reasoning remains in explicit CoT form to preserve accuracy. To learn this selective compression policy, SLT adopts a three-stage training strategy that combines span-level latent compression, reliability-aware future reasoning prediction, and trajectory-level reinforcement learning to optimize the trade-off between answer correctness and reasoning cost. Extensive experiments across four mathematical reasoning benchmarks demonstrate that SLT achieves 22.7% higher accuracy than latent reasoning baselines at comparable compression ratios, while reducing reasoning chain length by 58.4% with only 2.8% accuracy degradation compared to explicit CoT,Our code can be found in https://github.com/hunshi34/SLT.
Hui Xie, Jie Liu, Ziyue Qiao +1
Eindhoven University of Technology, Netherlands · School of Computing and Information Technology, Great Bay University, China
Chain-of-thought reasoning unfolds in discrete token space: each step is committed as text, errors propagate, and eliciting good traces presupposes traces to imitate. Reasoning instead in a model's continuous representation space - where intermediate states are vectors rather than words - sidesteps these constraints, but leaves open how those latent states should be computed. We approach this along two axes. First, we keep a large language model (LLM) frozen and use it for what it is already good at - modeling and decoding sequences - while a small auxiliary network supplies continuous latent thoughts as input. Second, we produce those latents by recurrence: a tiny recurrent reasoner refines them over many steps, decoupling the depth of computation from the size of the model, so that the latents are a product of iterative processing rather than a single forward pass. We instantiate this as Latent Recurrent Thoughts (LRT): a task-dedicated proposer supplies base latents, a recurrent reasoner refines them through bounded residual corrections, and the frozen LLM decodes the answer. On symbolic reasoning with answer supervision but no reasoning traces (Countdown-4, Sudoku) and on natural-language reasoning (HumanEval, MBPP, StrategyQA), LRT substantially outperforms prior frozen-decoder continuous-space reasoning methods under an identical decoder, prompt, data, and training budget, and outperforms non-thinking-mode chain-of-thought prompting on the same backbone at a small fraction of its inference compute.