Visual latent reasoning compresses rendered derivations into compact intermediate states, reducing textual reasoning overhead. Existing approaches differ in how they represent these states: continuous methods avoid vocabulary constraints, whereas discrete methods improve accuracy through quantization into a finite codebook. Our analysis of representative continuous and discrete systems identifies two functional requirements: answers must rely on latent states, and those states must carry valid, problem-specific reasoning. Continuous latents influence answers despite collapsed reasoning content, whereas discrete latents retain recoverable intermediate reasoning that answer prediction largely bypasses. To address these challenges, we propose Continuous Anchored Latent Reasoning (CALR), which connects latent formation with answer use through functional anchoring. With reference latents from information-balanced compression, CALR couples latent-mediated answer supervision with derivation-level semantic anchoring: the former routes answer supervision through intermediate states, while the latter grounds their decoded content in problem-specific derivations. A parallel-to-autoregressive curriculum develops sequential reasoning by conditioning subsequent latent blocks on generated prefixes. Evaluations on five mathematical reasoning benchmarks across model families show substantial accuracy gains. Under matched budgets, CALR gains 26.0 percentage points over a comparable continuous latent reasoning method. Further analyses show that its latents support answer prediction and carry problem-specific intermediate reasoning.
Figures & tables
Figure 1: The latent handoff challenge. (a) Continuous latents can support answer prediction without carrying problem-specific reasoning. (b) Discrete latents can encode recoverable reasoning yet be bypassed by the answer model. (c) CALR couples latent reliance (C1) with derivation anchoring for reasoning capability (C2).
Figure 2: Complementary handoff failures in continuous and discrete latents. (a) Both favor the question, with higher latent attention in the continuous system. (b) Content replacement has little effect ( Λcontent≈0 ); interface disruption yields accuracy drops of 0.42 and 0.015 , respectively.
Figure 3: CALR training framework: functional anchoring followed by parallel drafting and autoregressive refinement. Purple dashed paths denote the unified internal-reader route.
Resampling
GSM8K-Aug
Sim. (%) ↑
CE ↓
Global cross-attn.
92.8
0.183
Fixed anchors
95.3
0.158
Dynamic anchors
99.1
0.107
Table 5: Resampler ablations.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Parallel
Autoregressive
Reader question mask pr
1.0/0.5/0.3 by epoch
0.3
Gradient gate α
0→1 in epoch 2
0→1 over 3,000 steps
External weight w
1→0.5 in epoch 2
1→0.5 over 3,000 steps
Readiness checks
–
Every 500 steps; 200 problems
Reopening window
–
Steps 6,000–8,000
Appendix
Table 6: Unified-reader schedules. Epochs and steps are local to each phase.
Source
Training use
Examples
GSM8K-Aug
Resampler, reader, and producer
384,618
MATH
Resampler
11K
MathX-5M subset
Resampler
818K
Appendix
Table 7: Training data and their roles in CALR.
Hyperparameter
Stage I: Representation
Stage I: Readout
Stage II: Drafting
Stage II: Refinement
Training data
Rendered chains (Appendix B.2 )
GSM8K-Aug 385K
GSM8K-Aug 385K
GSM8K-Aug 385K
Optimizer / LR schedule
AdamW / cosine
AdamW / cosine
AdamW / cosine
AdamW / cosine
Peak learning rate
10−4
10−4
10−4
10−4
Warm-up steps
300
200
360
360
Steps (epochs)
12,000 ( ≈1 )
6,000 (2)
18,000 (3)
18,000 (3)
Global batch size
128
128
64
64
Appendix
Table 8: Per-stage training configurations. Rows marked “unified” apply only to unified CALR; the decoupled form uses the same shared settings with α=0 , w=1 , and no internal reader.
System
Question attention
Latent attention
Latent/question gradient ratio
RoT
0.51
0.33
1.12
DLR-SFT
0.72
0.05
0.13
Appendix
Table 9: Answer-stage attention and gradient measurements.
Condition
Metric
Set 1
Set 2
Feedback substitution + suffix regeneration
Expected
75.0
80.0
Reader-input substitution only
Expected
0.8
0.8
Same intermediate, different answer
Orig.
87.5
82.5
Appendix
Table 17: Intermediate-state substitution on two disjoint sets of 120 problems each. Expected denotes the percentage of outputs matching yexp ; Orig. denotes accuracy against the original answer y . Regeneration retains the original first two blocks at readout. All values are percentages.