Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST
Figures & tables
Figure 1: REST. The top row shows the four failures of Cross-Entropy (CE) only training of the thought. REST adds a β -weighted loss on to the CE, with one term per failure, under the same data and compute. (a) Accuracy gain for single- and multi-agent settings across Small (1–2B) and Large (3–4B) models. (b)–(f) REST thoughts spread apart instead of collapsing, reach a lower training CE and a higher accuracy, recover more accuracy when they replace the oracle text, reach a final answer more often, and keep their superposition.
Figure 2: Overview of REST. Latent systems pass a thought T through a trained outer link (left). REST adds a β -weighted loss built from the four properties of T to the CE objective (middle). CE-only thoughts collapse and REST improves accuracy under the same data and compute (right).
Figure 3: Why CE is not enough. Four failures of CE thoughts and the property that addresses each.
Figure 4: Single agent . Rin runs at each of the m′ steps, and Rψ once per round.
System
Role
Model
Light
Planner
Qwen3-1.7B
Refiner
Llama-3.2-1B-Instruct
Solver
Qwen2.5-Math-1.5B-Instruct
Scaled
Planner
Gemma-3-4B-it
Refiner
Llama-3.2-3B-Instruct
Solver
Qwen3.5-4B
Table 1: Light and Scaled systems.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.6
80.9
27.8
65.6
12.2
72.2
27.3
63.3
27.1
79.2
27.7
35.5
Base
Base
CE only
Token
557
862
905
7226
1016
6967
911
1709
1177
739
477
1289
Base
Base
REST (ours), single property
Acc.
72.4
80.4
23.3
78.9
16.7
83.3
23.9
63.5
29.7
79.4
32.9
40.2
↑1.0
↑4.8
Causality
Token
551
1008
907
8677
1008
8165
851
2216
1113
955
570
1662
−0.9%
+20.7%
Table 2: Single-agent, Light vs Scaled, at r=1 and r=3 respectively, each row at its best β .
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
86.6
22.2
80.0
17.8
60.0
24.9
59.1
28.9
79.7
32.5
33.5
Base
Base
CE only
Token
550
1252
849
10485
884
5420
749
2572
868
829
604
1029
Base
Base
REST (ours), single property
Acc.
77.8
86.8
27.8
80.0
22.2
86.7
28.6
54.9
30.0
78.7
33.9
39.1
↑3.8
↑4.6
Causality
Token
601
1158
894
8576
978
8278
888
1568
1049
996
643
1924
+12.2%
+4.2%
Table 3: Multi-agent, Light vs Scaled, at r=1 and r=3 respectively, each row at its best β .
Figure 5: Acc. Change vs r
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
87.7
22.2
76.7
17.8
86.7
24.9
63.1
28.9
80.2
32.5
39.0
Base
Base
CE only
Token
550
1018
849
9009
884
8488
749
2379
868
1137
604
1507
Base
Base
Acc.
76.9
85.9
30.0
83.3
22.2
83.3
28.8
61.8
27.3
80.3
32.0
39.5
↑3.3
↑0.1
CODI ( β=20 )
Token
598
1035
933
9344
987
9153
861
2293
1101
1097
682
1602
+14.6%
+4.2%
Acc.
76.7
82.0
27.8
77.8
21.1
78.9
26.9
54.5
28.9
79.0
34.9
35.0
↑3.2
↓4.4
Table 4: REST against auxiliary-loss baselines, Light vs Scaled.
Figure 6: REST against CE-only. (a) Decoded thoughts compared to output / input. (b) Replacing oracle text with thought. (c) Answer rate vs token usage. (d) Effective Superposition.
Figure 7: REST against CE-only. (a) PCA for T . CE thoughts collapse into dense clusters, while REST spreads thoughts apart. (b) Training CE under causality. More details in Appendix D .
Figure 8: Positioning of REST.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Light
Scaled
MATH500
1000
2000
AIME2025
8192
16000
AIME2026
8192
16000
GPQA-D
4000
4000
MedQA
4000
4000
MBPP+
4000
4000
Appendix
Table 5: Generation budget.
Model
Metric
MATH500
GPQA-D
MedQA
AIME25
AIME26
MBPP+
LCB-v6
Qwen3-1.7B
Acc.
68.4
33.5
46.2
20.0
22.2
57.3
22.8
Token
546
1009
472
1962
2695
80
608
Llama-3.2-1B
Acc.
25.0
27.4
37.0
1.1
2.2
41.7
4.3
Token
417
629
369
2412
2225
309
288
Qwen2.5-Math-1.5B
Acc.
76.3
29.6
28.4
27.8
22.2
33.1
2.9
Token
530
807
1061
908
987
693
669
Appendix
Table 6: Frozen base LLMs, no system.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.6
81.1
27.8
80.0
12.2
80.0
27.3
61.8
27.1
79.8
27.7
37.3
Base
Base
CE only
Token
557
880
905
8272
1016
8468
911
2081
1177
767
477
1353
Base
Base
REST (ours), single property
Acc.
72.4
83.3
23.3
68.9
16.7
76.7
23.9
62.8
29.7
79.9
32.9
38.1
↑1.0
↓1.7
Causality ( β=0.3 )
Token
551
842
907
8335
1008
7292
851
1999
1113
739
570
1369
−0.9%
−5.7%
Appendix
Table 7: Single-agent (round r=1 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.9
80.9
26.7
65.6
20.0
72.2
26.9
63.3
27.8
79.2
28.1
35.5
Base
Base
CE only
Token
613
862
912
7226
1020
6967
941
1709
1312
739
667
1289
Base
Base
REST (ours), single property
Acc.
74.9
80.4
26.7
78.9
18.9
83.3
29.1
63.5
27.6
79.4
30.0
40.2
↑1.1
↑4.8
Causality ( β=0.3 )
Token
619
1008
910
8677
952
8165
899
2216
1112
955
654
1662
−5.8%
+20.7%
Appendix
Table 8: Single-agent (round r=3 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
87.7
22.2
76.7
17.8
86.7
24.9
63.1
28.9
80.2
32.5
39.0
Base
Base
CE only
Token
550
1018
849
9009
884
8488
749
2379
868
1137
604
1507
Base
Base
REST (ours), single property
Acc.
77.8
86.4
27.8
80.0
22.2
85.6
28.6
63.1
30.0
82.7
33.9
39.5
↑3.8
↑0.7
Causality
Token
601
1041
894
9538
978
9430
888
2285
1049
1115
643
1545
+12.2%
+6.0%
Appendix
Table 9: Multi-agent (round r=1 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
70.9
86.6
22.2
80.0
15.6
60.0
27.3
59.1
28.9
79.7
30.3
33.5
Base
Base
CE only
Token
716
1252
796
10485
891
5420
973
2572
1145
829
857
1029
Base
Base
REST (ours), single property
Acc.
75.4
86.8
30.0
80.0
14.4
86.7
28.3
54.9
28.1
78.7
34.1
39.1
↑2.5
↑4.6
Causality
Token
792
1158
933
8576
1024
8278
1087
1568
1304
996
777
1924
+10.0%
+4.2%
Appendix
Table 10: Multi-agent (round r=3 ), Light vs Scaled.
Method
Metric
Math500
AIME2025
AIME2026
GPQA-D
MedQA
Code Gen.
Avg. Change
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Light
Scaled
Acc.
71.1
87.7
22.2
76.7
17.8
86.7
24.9
63.1
28.9
80.2
32.5
39.0
Base
Base
CE only
Token
550
1018
849
9009
884
8488
749
2379
868
1137
604
1507
Base
Base
REST (ours), causality
Acc.
76.1
86.6
31.1
86.7
20.0
83.3
26.6
65.2
28.2
83.3
33.3
41.5
↑3.0
↑2.2
β=0.1
Token
608
999
940
10437
1009
10161
941
2187
1151
1066
613
1882
+16.9%
+13.6%
Appendix
Table 11: Full property sweep, multi-agent (round r=1 ), Light vs Scaled.
Figure 10: Decoded thoughts compared to the producer’s output.
Figure 11: PCA for T . CE thoughts collapse into dense clusters, where each dashed ring holds all questions of one training run, while REST spreads thoughts apart.
Figure 12: Training CE. (a) Training CE loss for causality at β=3.0 . (b) Mean of last CE 200 steps relative to CE-only against Acc. Change. Bars span seeds.
Figure 13: Accuracy against the latent budget m′ for the Scaled system at r=1 .