Continuous reasoning has emerged as a promising way to improve reasoning in large language models (LLMs). Yet we still lack a clear principle for deciding what a latent state should preserve. Reasoning by superposition shows that a single latent state can encode several search alternatives and expand them in parallel. We ask how those states should be weighted as reasoning proceeds. A natural choice is to preserve only the states active at the frontier step, since keeping every reached state appears to spread a limited hidden width too thin. We show that the opposite can hold. When later computation draws on several reached states, a cumulative state can guide attention correctly at a smaller hidden width than a frontier state that stores fewer states. At the same width, the cumulative state therefore keeps more intermediate states available for later reasoning. More generally, equal cumulative weights are optimal when future queries are unknown and remain close to the best task-specific weights when those queries are known. Experiments with two-layer and GPT-2 Transformers reproduce the predicted width advantage and show that unequal weights fail first on the states that receive the least weight. This suggests a important principle: keep reached states equally weighted, and restore equal weights as computation proceeds.
Figures & tables
Figure 1: The cumulative-state hypothesis from latent state to computation. (a) The frontier state keeps only the states active now, whereas the cumulative state keeps all reached states for the same final task. (b) Each update adds new states; equalizing the weights keeps every reached state available. (c) Equal cumulative weights avoid a weakly represented state and are optimal when future use is unknown.
Figure 2: Why remembering more can require less width. (a) Across many ways of using the reached history state, cumulative state stays below the predicted width bound, while the deliberately hardest examples approach it. (b) As the reached history grows, cumulative keeps useful signal for the final task, whereas frontier discards earlier states and becomes harder to use.
Figure 3: Trained Transformers reproduce the predicted width advantage. The shared legend shows six ways to weight the same reached history. (a1) Results for the two-layer block model; (a2) results for the GPT-2-block model, with bands showing variation across runs. Every curve in a panel uses the same trained attention module, and equal cumulative weighting reaches 90% accuracy on the hardest tested task first. (b) Five paired runs confirm that cumulative reaches the same held-out accuracy at a smaller width than frontier in both architectures. Here d0.9 means the smallest interpolated width reaching 90% accuracy; Ideal supplies the intended state directly.
Figure 4: Why equal cumulative weights are safest. (a) Each point compares the width needed to keep states usable during reasoning and the width needed to use them for the final answer; lower is better. (b) The heatmap shows how much state is assigned to each position in the reached history. (c) Using the same trained attention module, states with small weights are the first to be missed. The summary reports success on the whole routing step and the rate at which inactive states are selected by mistake.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Split
# traces
∣V∣
∣E∣
solution length
Train
14,785
22.8
36.5
3.5
Validation
257
22.7
36.3
3.5
Test
419
22.7
36.0
3.5
Appendix
Table 1: ProsQA graph-4 statistics. Graph sizes and solution lengths are averaged over traces.
Split
Used traces
Queries
Support fraction
Realized Cmax
Train
14,785
118,280
[0.750,0.857]
1.319
Validation
120
480
[0.750,0.857]
1.271
Test
419
3,352
[0.750,0.857]
1.287
Appendix
Table 2: Audit of the fixed nonlocalized query bank. “Queries” counts coefficient-space queries before repeated Gaussian test draws.
Component
Configuration
Widths
16,24,32,40,48,56,64,80,96,112,128
Training seeds / embedding seeds
0:4 / 5100:5104
Query-bank size per trace (train/val/test)
8/4/8
Decoys (train/val/test)
256/1,024/4,096
Gaussian draws (val/test)
1/4
Epochs and optimizer
4; AdamW, learning rate 3×10−3
Appendix
Table 3: Configuration of the lightweight two-block paired experiment.
d
Learned frontier
Learned cumulative
Ideal frontier
Ideal cumulative
16
0.048±0.001
0.254±0.012
0.052±0.001
0.442±0.003
24
0.114±0.004
0.574±0.008
0.120±0.005
0.790±0.004
32
0.210±0.007
0.793±0.015
0.216±0.007
0.947±0.005
40
0.322±0.005
0.899±0.010
0.328±0.006
0.990±0.002
48
0.430±0.008
0.946±0.004
0.436±0.008
0.9987±0.0002
56
0.533±0.010
0.972±0.005
0.540±0.009
0.9999±0.0002
Appendix
Table 4: Held-out top-1 accuracy at every tested width.
Seed
Frontier
Cumulative
C/F
0
108.9
39.2
0.360
1
112.1
39.9
0.356
2
110.3
40.4
0.366
3
111.7
41.6
0.372
4
110.1
40.6
0.369
Mean
110.6
40.3
0.365
Appendix
Table 5: Interpolated empirical critical dimension d0.9 for every lightweight-model training seed.
Component
Configuration
Backbone
causal GPT-2; 2 layers, 8 heads, FFN width 4d
Transformer details
GELU, context 512, dropout 0.1, residuals and LayerNorm
Tested widths / maximum codebook width
32,64,96,128,192,256,384 / 768
Model seeds / embedding seeds
0:4 / 6100:6104
Used traces (train/validation/test)
800/120/120
Queries per trace (train/validation/test)
8/4/4
Appendix
Table 6: Complete configuration of the paired GPT-2 width sweep.
Table 8: GPT-2 interpolated critical dimension d0.9 for every model seed.
Codebook seed
Uniform cumulative
Recency-heavy
Random mild
4903
0.896/0.032
0.576/0.572
0.533/1.318
5004
0.904/0.028
0.594/0.585
0.541/1.282
5105
0.907/0.028
0.593/0.557
0.536/1.312
5206
0.913/0.026
0.587/0.564
0.541/1.336
5307
0.909/0.029
0.578/0.570
0.534/1.280
Mean
0.906/0.029
0.586/0.570
0.537/1.306
Appendix
Table 9: Raw reusable-token routing results at d=192 . Each cell reports mean joint route-step success (fraction) / per-inactive-key false-positive rate (%). A route step succeeds iff every currently reached node is accepted and every unreached graph node and all 512 Gaussian decoys are rejected.
Current transformers discard their rich latent residual stream between positions, reconstructing latent reasoning context at each new position and leaving potential reasoning capacity untapped. The State Stream Transformer (SST) V2 enables parameter-efficient reasoning in continuous latent space through an FFN-driven nonlinear recurrence at each decoder layer, where latent states are streamed horizontally across the full sequence via a learned blend. This same mechanism supports continuous latent deliberation per position at inference time, dedicating additional FLOPs to exploring abstract reasoning before committing to a token. A two-pass parallel training procedure resolves the sequential dependency of the recurrence to allow compute-efficient training. Hidden state analysis shows the state stream facilitates reasoning through exploration of distinct semantic basins in continuous latent space, where transitions at content-dependent positions move the model into a substantially different Bayesian posterior, directly influencing the latent space at future positions. We also find, via a learned probe, that at the first generated token position, the latent state already predicts whether the eventual answer will survive or break under additional latent computation for every subsequent position. Co-trained into an existing 27B backbone using only a small dataset of GSM8K examples, the SST delivers a +15.15 point gain over a fine-tuning-matched baseline on out-of-distribution GPQA-Diamond and cuts that same baseline's remaining GSM8K errors by 46%, together showing that the reasoning improvement is attributable to the architectural mechanism rather than scale or training data. On GPQA-Diamond, the resulting 27B SST also achieves higher accuracy than several larger open-weight and proprietary systems, including open-weight models up to 25 times larger.
Reasoning in large language models is often discussed as a single capability, but some of its gains may stem from simpler underlying operations. We examine two such primitives, recall and state-tracking, through five controlled task families centered on state-based recall, and compare matched transformer and hybrid architectures with and without reasoning augmentation. Across the suite, reasoning-augmented variants substantially outperform instruction-only variants, often by large margins. This pattern is consistent with the State over Tokens view: externalized reasoning traces help because they carry the intermediate state forward in token space. By contrast, hybrid inductive bias does not yield a uniform advantage in accuracy once reasoning tokens are available. When architectural differences do appear, they follow task structure: the hybrid Think model is more robust on strictly sequential chained updates, whereas the transformer Think model is more robust on flat multi-hop retrieval. We therefore cast the main contribution of this study as a descriptive account of what drives performance on state-based recall tasks: reasoning-token augmentation appears to be the dominant factor, while hybrid advantages are narrower, task-dependent, and potentially more about inference efficiency than overall capability. We also release the codebase and data required to reproduce these results.
Shivam Rawat, Lucie Flek, Florian Mai +1
Lamarr Institute for Machine Learning and Artificial Intelligence · Rheinische Friedrich-Wilhelms-Universität Bonn
Large Language Models (LLMs) increasingly rely on intermediate reasoning, yet explicit Chain-of-Thought (CoT) suffers from a linguistic space bottleneck: each thought must be decoded into tokens, causing high inference overhead. Latent reasoning moves deliberation into continuous space, but existing methods mostly learn deterministic or reward-maximizing paths, lacking a principled way to allocate probability across trajectories with different correctness and costs. We propose Latent Thought Flow (LTF), which models reasoning as variable-length continuous trajectories and trains a sampler to match a reward-induced posterior over answer quality and computation cost. We instantiate this with a continuous GFlowNet using stochastic latent transitions. To handle sparse answer supervision, we introduce an Entropy-Weighted Subtrajectory Balance objective for intermediate rewards and a reference-prior regularizer to anchor exploration. Experiments under finetuning and transfer learning settings show that LTF outperforms explicit CoT and latent reasoning baselines, improving accuracy by 9.5% while reducing reasoning length by 27.2% on average compared with strong latent reasoning baselines.