Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST
Figures & tables
Figure 1: REST. The top row shows the four failures of Cross-Entropy (CE) only training of the thought. REST adds a β -weighted loss on to the CE, with one term per failure, under the same data and compute. (a) Accuracy gain for single- and multi-agent settings across Small (1–2B) and Large (3–4B) models. (b)–(f) REST thoughts spread apart instead of collapsing, reach a lower training CE and a higher accuracy, recover more accuracy when they replace the oracle text, reach a final answer more often, and keep their superposition.
Figure 2: Overview of REST. Latent systems pass a thought T through a trained outer link (left). REST adds a β -weighted loss built from the four properties of T to the CE objective (middle). CE-only thoughts collapse and REST improves accuracy under the same data and compute (right).
Figure 3: Why CE is not enough. Four failures of CE thoughts and the property that addresses each.
Figure 4: Single agent . Rin runs at each of the m′ steps, and Rψ once per round.
System
Role
Model
Light
Planner
Qwen3-1.7B
Refiner
Llama-3.2-1B-Instruct
Solver
Qwen2.5-Math-1.5B-Instruct
Scaled
Planner
Gemma-3-4B-it
Refiner
Llama-3.2-3B-Instruct
Solver
Qwen3.5-4B
Table 1: Light and Scaled systems.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.6
80.9
27.8
65.6
12.2
72.2
27.3
63.3
27.1
79.2
27.7
35.5
Base
Base
CE only
Token
557
862
905
7226
1016
6967
911
1709
1177
739
477
1289
Base
Base
REST (ours), single property
Acc.
72.4
80.4
23.3
78.9
16.7
83.3
23.9
63.5
29.7
79.4
32.9
40.2
↑1.0
↑4.8
Causality
Token
551
1008
907
8677
1008
8165
851
2216
1113
955
570
1662
−0.9%
+20.7%
Table 2: Single-agent, Light vs Scaled, at r=1 and r=3 respectively, each row at its best β .
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
86.6
22.2
80.0
17.8
60.0
24.9
59.1
28.9
79.7
32.5
33.5
Base
Base
CE only
Token
550
1252
849
10485
884
5420
749
2572
868
829
604
1029
Base
Base
REST (ours), single property
Acc.
77.8
86.8
27.8
80.0
22.2
86.7
28.6
54.9
30.0
78.7
33.9
39.1
↑3.8
↑4.6
Causality
Token
601
1158
894
8576
978
8278
888
1568
1049
996
643
1924
+12.2%
+4.2%
Table 3: Multi-agent, Light vs Scaled, at r=1 and r=3 respectively, each row at its best β .
Figure 5: Acc. Change vs r
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
87.7
22.2
76.7
17.8
86.7
24.9
63.1
28.9
80.2
32.5
39.0
Base
Base
CE only
Token
550
1018
849
9009
884
8488
749
2379
868
1137
604
1507
Base
Base
Acc.
76.9
85.9
30.0
83.3
22.2
83.3
28.8
61.8
27.3
80.3
32.0
39.5
↑3.3
↑0.1
CODI ( β=20 )
Token
598
1035
933
9344
987
9153
861
2293
1101
1097
682
1602
+14.6%
+4.2%
Acc.
76.7
82.0
27.8
77.8
21.1
78.9
26.9
54.5
28.9
79.0
34.9
35.0
↑3.2
↓4.4
Table 4: REST against auxiliary-loss baselines, Light vs Scaled.
Figure 6: REST against CE-only. (a) Decoded thoughts compared to output / input. (b) Replacing oracle text with thought. (c) Answer rate vs token usage. (d) Effective Superposition.
Figure 7: REST against CE-only. (a) PCA for T . CE thoughts collapse into dense clusters, while REST spreads thoughts apart. (b) Training CE under causality. More details in Appendix D .
Figure 8: Positioning of REST.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Light
Scaled
MATH500
1000
2000
AIME2025
8192
16000
AIME2026
8192
16000
GPQA-D
4000
4000
MedQA
4000
4000
MBPP+
4000
4000
Appendix
Table 5: Generation budget.
Model
Metric
MATH500
GPQA-D
MedQA
AIME25
AIME26
MBPP+
LCB-v6
Qwen3-1.7B
Acc.
68.4
33.5
46.2
20.0
22.2
57.3
22.8
Token
546
1009
472
1962
2695
80
608
Llama-3.2-1B
Acc.
25.0
27.4
37.0
1.1
2.2
41.7
4.3
Token
417
629
369
2412
2225
309
288
Qwen2.5-Math-1.5B
Acc.
76.3
29.6
28.4
27.8
22.2
33.1
2.9
Token
530
807
1061
908
987
693
669
Appendix
Table 6: Frozen base LLMs, no system.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.6
81.1
27.8
80.0
12.2
80.0
27.3
61.8
27.1
79.8
27.7
37.3
Base
Base
CE only
Token
557
880
905
8272
1016
8468
911
2081
1177
767
477
1353
Base
Base
REST (ours), single property
Acc.
72.4
83.3
23.3
68.9
16.7
76.7
23.9
62.8
29.7
79.9
32.9
38.1
↑1.0
↓1.7
Causality ( β=0.3 )
Token
551
842
907
8335
1008
7292
851
1999
1113
739
570
1369
−0.9%
−5.7%
Appendix
Table 7: Single-agent (round r=1 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.9
80.9
26.7
65.6
20.0
72.2
26.9
63.3
27.8
79.2
28.1
35.5
Base
Base
CE only
Token
613
862
912
7226
1020
6967
941
1709
1312
739
667
1289
Base
Base
REST (ours), single property
Acc.
74.9
80.4
26.7
78.9
18.9
83.3
29.1
63.5
27.6
79.4
30.0
40.2
↑1.1
↑4.8
Causality ( β=0.3 )
Token
619
1008
910
8677
952
8165
899
2216
1112
955
654
1662
−5.8%
+20.7%
Appendix
Table 8: Single-agent (round r=3 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
87.7
22.2
76.7
17.8
86.7
24.9
63.1
28.9
80.2
32.5
39.0
Base
Base
CE only
Token
550
1018
849
9009
884
8488
749
2379
868
1137
604
1507
Base
Base
REST (ours), single property
Acc.
77.8
86.4
27.8
80.0
22.2
85.6
28.6
63.1
30.0
82.7
33.9
39.5
↑3.8
↑0.7
Causality
Token
601
1041
894
9538
978
9430
888
2285
1049
1115
643
1545
+12.2%
+6.0%
Appendix
Table 9: Multi-agent (round r=1 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.9
86.6
22.2
80.0
15.6
60.0
27.3
59.1
28.9
79.7
30.3
33.5
Base
Base
CE only
Token
716
1252
796
10485
891
5420
973
2572
1145
829
857
1029
Base
Base
REST (ours), single property
Acc.
75.4
86.8
30.0
80.0
14.4
86.7
28.3
54.9
28.1
78.7
34.1
39.1
↑2.5
↑4.6
Causality
Token
792
1158
933
8576
1024
8278
1087
1568
1304
996
777
1924
+10.0%
+4.2%
Appendix
Table 10: Multi-agent (round r=3 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
87.7
22.2
76.7
17.8
86.7
24.9
63.1
28.9
80.2
32.5
39.0
Base
Base
CE only
Token
550
1018
849
9009
884
8488
749
2379
868
1137
604
1507
Base
Base
REST (ours), causality
Acc.
76.1
86.6
31.1
86.7
20.0
83.3
26.6
65.2
28.2
83.3
33.3
41.5
↑3.0
↑2.2
β=0.1
Token
608
999
940
10437
1009
10161
941
2187
1151
1066
613
1882
+16.9%
+13.6%
Appendix
Table 11: Full property sweep, multi-agent (round r=1 ), Light vs Scaled.
Figure 10: Decoded thoughts compared to the producer’s output.
Figure 11: PCA for T . CE thoughts collapse into dense clusters, where each dashed ring holds all questions of one training run, while REST spreads thoughts apart.
Figure 12: Training CE. (a) Training CE loss for causality at β=3.0 . (b) Mean of last CE 200 steps relative to CE-only against Acc. Change. Bars span seeds.
Figure 13: Accuracy against the latent budget m′ for the Scaled system at r=1 .
Latent reasoning allows language models to carry out intermediate reasoning in continuous latent representations rather than fully externalizing it as discrete chains of thought. However, assigning credit to such latent thoughts from answer-only rewards is difficult: a single final answer mixes thought quality with answer-sampling noise. We propose \textbf{Latent Thought Credit (LTC)}, a hierarchical credit-assignment framework for latent reasoning. For each prompt, LTC samples multiple latent thoughts, fixes the context after each thought, and estimates thought-level expected reward by averaging rewards over multiple answers generated from that fixed context. LTC uses thought-level advantages to optimize the latent-thought phase, answer-level advantages to optimize the answer phase, and an advantage-weighted thought-matching objective that helps the policy reproduce high-credit latent thoughts. We instantiate LTC in a GRPO-style on-policy training framework and evaluate it across mathematical reasoning and STEM multiple-choice tasks. LTC achieves the best average accuracy among the compared methods, while ablations and fixed-context diagnostics show that multi-answer estimation reduces reward-estimation error and mitigates ambiguous or incorrect thought-level credit.
Xuyang Zhao, Liting Zhang, Zichen Xu +4
TMCC, College of Computer Science, Nankai University, Tianjin, China · Lingxi (Beijing) Technology Co., Ltd.
Chain-of-thought reasoning unfolds in discrete token space: each step is committed as text, errors propagate, and eliciting good traces presupposes traces to imitate. Reasoning instead in a model's continuous representation space - where intermediate states are vectors rather than words - sidesteps these constraints, but leaves open how those latent states should be computed. We approach this along two axes. First, we keep a large language model (LLM) frozen and use it for what it is already good at - modeling and decoding sequences - while a small auxiliary network supplies continuous latent thoughts as input. Second, we produce those latents by recurrence: a tiny recurrent reasoner refines them over many steps, decoupling the depth of computation from the size of the model, so that the latents are a product of iterative processing rather than a single forward pass. We instantiate this as Latent Recurrent Thoughts (LRT): a task-dedicated proposer supplies base latents, a recurrent reasoner refines them through bounded residual corrections, and the frozen LLM decodes the answer. On symbolic reasoning with answer supervision but no reasoning traces (Countdown-4, Sudoku) and on natural-language reasoning (HumanEval, MBPP, StrategyQA), LRT substantially outperforms prior frozen-decoder continuous-space reasoning methods under an identical decoder, prompt, data, and training budget, and outperforms non-thinking-mode chain-of-thought prompting on the same backbone at a small fraction of its inference compute.
Large Language Models (LLMs) increasingly rely on intermediate reasoning, yet explicit Chain-of-Thought (CoT) suffers from a linguistic space bottleneck: each thought must be decoded into tokens, causing high inference overhead. Latent reasoning moves deliberation into continuous space, but existing methods mostly learn deterministic or reward-maximizing paths, lacking a principled way to allocate probability across trajectories with different correctness and costs. We propose Latent Thought Flow (LTF), which models reasoning as variable-length continuous trajectories and trains a sampler to match a reward-induced posterior over answer quality and computation cost. We instantiate this with a continuous GFlowNet using stochastic latent transitions. To handle sparse answer supervision, we introduce an Entropy-Weighted Subtrajectory Balance objective for intermediate rewards and a reference-prior regularizer to anchor exploration. Experiments under finetuning and transfer learning settings show that LTF outperforms explicit CoT and latent reasoning baselines, improving accuracy by 9.5% while reducing reasoning length by 27.2% on average compared with strong latent reasoning baselines.