We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer. It iteratively plans a horizon of candidate intermediate questions, selects one to answer, and replans from the updated history. ReHoPER is task-agnostic, using the same generic instructions across datasets and models without labeled data or task-specific prompt design. Across multiple datasets, including iLLC, a new controlled benchmark for compositional reasoning, ReHoPER outperforms strong baselines, with the largest gains in the most compositional settings. Our implementation and the iLLC generator are publicly available to support future work.
Figures & tables
Figure 1 : ReHoPER overview. Layered frames indicate independent paths. Each path plans questions, answers one, and replans from updated history. A majority vote over the initial and per-step predictions across paths yields the final prediction.
MuSR MM
MuSR OP
MuSR TA
IO
CoT
PS
IO
CoT
PS
IO
CoT
PS
Model
SD
SC
RH
SC
RH
SC
RH
SD
SC
RH
SC
RH
SC
RH
SD
SC
RH
SC
RH
SC
RH
Gemma 3
57.6
54.4
61.2
61.2
62.0
61.6
65.6
50.0
50.4
55.1
51.6
55.5
54.7
55.9
50.0
58.8
47.6
56.0
51.2
58.0
56.0
gpt-oss
74.8
70.8
76.0
70.4
76.4
73.2
76.8
53.9
55.5
53.9
53.1
53.5
57.0
53.5
70.8
69.2
68.0
56.4
63.2
65.2
68.0
Llama 3.1 8B
54.0
56.8
62.4
56.4
60.8
57.2
63.6
44.5
49.2
53.1
52.0
48.4
52.0
48.4
34.8
35.2
43.6
43.2
47.6
45.2
47.6
Llama 3.1 70B
64.8
65.6
71.2
66.0
70.4
66.8
71.2
42.6
46.1
55.9
46.1
54.3
48.4
55.9
57.6
62.0
57.6
62.0
59.6
60.4
59.2
Table 1 : Accuracy (%) of ReHoPER (RH) on MuSR, MoreHopQA, and iLLC benchmarks under IO, CoT, and PS prefixes. SD refers to the Self-Discover baseline. SC denotes simple question answering with Self-Consistency using the corresponding output prefix, while RH applies our planning approach with that prefix. Green , red , and gray indicate performance changes of ReHoPER relative to SC. The bottom-right block reports averages over the eight displayed subsets for each model and prompt.
MuSR MM
MoreHopQA
IO
CoT
PS
IO
CoT
PS
Model
-RP
-LA
RH
-RP
-LA
RH
-RP
-LA
RH
-RP
-LA
RH
-RP
-LA
RH
-RP
-LA
RH
Gemma 3
57.6
58.8
61.2
59.2
58.4
62.0
62.4
62.4
65.6
64.0
66.0
64.7
66.7
66.7
68.7
70.0
69.3
74.0
gpt-oss
72.4
70.0
76.0
72.0
68.0
76.4
74.0
71.2
76.8
80.7
78.7
78.7
79.3
79.3
80.0
80.7
78.7
81.3
Llama 3.1 8B
61.6
58.4
62.4
61.6
61.2
60.8
64.8
60.4
63.6
52.0
57.3
57.3
56.7
58.7
60.7
60.7
66.0
63.3
Llama 3.1 70B
70.4
70.0
71.2
72.0
72.4
70.4
71.6
71.2
71.2
68.7
74.0
71.3
70.7
72.7
71.3
72.7
74.0
74.0
Table 2 : Planning ablations without replanning (-RP), without look-ahead (-LA), and the full ReHoPER method (RH) across three prompting styles, IO, CoT, and PS, on MuSR MM, MoreHopQA, and iLLC L2-4. Values are accuracies (%), and green and light green mark the best and second best variant, respectively, for each (model, dataset, prefix) triplet. The bottom-right block reports averages over all datasets for each model and prompt.
Figure 2 : Answer aggregation for ReHoPER+PS on iLLC L2-4 across models. Top: accuracy over reasoning steps; bottom: accuracy over stochastic paths. CV is cumulative voting, CSV voting over each path’s latest prediction, and CSM their mean correctness. CPV/CPM are voting/mean correctness within a path; APV/APM are their averages across paths.
Table 5
iLLC L1-4
iLLC L1-6
iLLC L2-4
iLLC L2-6
Model
SC 61
SC 181
RH
SC 61
SC 181
RH
SC 61
SC 181
RH
SC 61
SC 181
RH
Llama 3.1 70B
Runtime
1:16
3:37
2:42
1:26
4:23
4:12
2:11
6:23
7:56
2:37
9:04
11:15
Accuracy
92.1
92.4
92.7
76.0
76.9
80.1
32.5
35.1
79.3
17.2
16.7
68.5
Mistral Small
Runtime
0:38
1:38
0:54
0:43
2:03
0:55
0:48
2:08
1:15
0:56
2:41
1:13
Accuracy
90.7
90.8
91.3
76.1
76.9
73.5
3.5
2.4
11.6
0.0
0.1
2.0
Phi-4
Runtime
0:45
1:46
1:06
0:52
2:31
1:57
0:50
2:01
1:18
1:00
2:55
2:21
Table 4 : Runtime (h:mm) and accuracy (%) comparison between Self-Consistency (SC 61 ) and ReHoPER (RH) across iLLC L1 and L2 subsets. SC 181 uses 181 samples as a higher-compute SC comparison.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
MuSR
iLLC
IO
CoT
PS
IO
CoT
PS
Model
SD
SC
RH
SC
RH
SC
RH
SD
SC
RH
SC
RH
SC
RH
Gemma 3
52.5
54.5
54.6
56.3
56.2
58.1
59.2
24.7
7.8
16.6
30.5
30.5
37.1
32.5
gpt-oss
66.5
65.2
66.0
60.0
64.4
65.1
66.1
62.2
85.9
83.9
78.3
83.4
52.7
83.9
Llama 3.1 8B
44.4
47.1
53.0
50.5
52.3
51.5
53.2
12.2
28.3
26.9
26.5
31.7
35.0
37.1
Llama 3.1 70B
55.0
57.9
61.6
58.0
61.4
58.5
62.1
39.9
40.3
69.9
44.5
69.9
56.6
77.3
Appendix
Table 5: Macro-aggregated accuracy (%) on three benchmark families. MuSR averages MM, OP, and TA; iLLC averages L1-4, L1-6, L1-8, L2-4, L2-6, and L2-8; and MoreHopQA is reported separately. The final block is an equal-weight average of these three family-level results. SD denotes Self-Discover. Green , red , and gray indicate improvements, decreases, and ties of ReHoPER (RH) relative to Self-Consistency (SC), respectively.
Figure 4 : Step-wise aggregation analysis for ReHoPER under the IO answer format. Each panel shows accuracy as a function of the reasoning step for one model–dataset pair. CV denotes cumulative voting over all predictions up to the current step; CSV votes over one carried-forward prediction per path at the current step; CSM reports the mean carried-forward per-path correctness. Shaded bands show the standard deviation for the per-path CSM diagnostic.
Figure 5 : Step-wise aggregation analysis for ReHoPER under the CoT answer format. Each panel shows accuracy as a function of the reasoning step for one model–dataset pair. CV denotes cumulative voting over all predictions up to the current step; CSV votes over one carried-forward prediction per path at the current step; CSM reports the mean carried-forward per-path correctness. Shaded bands show the standard deviation for the per-path CSM diagnostic.
Figure 6 : Step-wise aggregation analysis for ReHoPER under the PS answer format. Each panel shows accuracy as a function of the reasoning step for one model–dataset pair. CV denotes cumulative voting over all predictions up to the current step; CSV votes over one carried-forward prediction per path at the current step; CSM reports the mean carried-forward per-path correctness. Shaded bands show the standard deviation for the per-path CSM diagnostic.
Figure 7 : Path-wise aggregation analysis for ReHoPER under the IO answer format. CV cumulatively votes over predictions from paths 1 through j , including the shared initial prediction only once. CPV and CPM summarize the current path alone, while APV and APM show the average path-local vote and mean correctness across all paths.
Figure 8 : Path-wise aggregation analysis for ReHoPER under the CoT answer format. CV cumulatively votes over predictions from paths 1 through j , including the shared initial prediction only once. CPV and CPM summarize the current path alone, while APV and APM show the average path-local vote and mean correctness across all paths.
Figure 9 : Path-wise aggregation analysis for ReHoPER under the PS answer format. CV cumulatively votes over predictions from paths 1 through j , including the shared initial prediction only once. CPV and CPM summarize the current path alone, while APV and APM show the average path-local vote and mean correctness across all paths.
Figure 10 : Accuracy gains from allocating iterations to steps vs. paths under the IO answer format.
Figure 11 : Accuracy gains from allocating iterations to steps vs. paths under the CoT answer format.
Figure 12 : Accuracy gains from allocating iterations to steps vs. paths under the PS answer format.
MuSR MM
Model
#IQ
#IQ/Path std
#uniqIQ
#uniqIQ/Path std
%uniqIQ
%uniqIQ/Path
Gemma 3
60.00
10.000.00
59.29
10.000.00
98.82%
100.00%
gpt-oss
39.38
6.563.72
39.29
6.563.72
99.77%
100.00%
Llama 3.1 8B
59.95
9.990.02
59.54
9.990.02
99.32%
100.00%
Llama 3.1 70B
60.00
10.000.00
56.72
10.000.00
94.53%
100.00%
Mistral Small
59.96
9.990.01
55.86
9.990.01
93.16%
100.00%
Appendix
Table 6 : Statistics of generated intermediate questions (IQs) in ReHoPER on MuSR and MoreHopQA . We report the total number of generated IQs, the average number of IQs per path with standard deviation, the number of unique IQs, the average number of unique IQs per path with standard deviation, and the corresponding uniqueness percentages.
iLLC L1-4
Model
#IQ
#IQ/Path std
#uniqIQ
#uniqIQ/Path std
%uniqIQ
%uniqIQ/Path
Gemma 3
34.98
5.830.90
17.23
5.830.90
49.25%
100.00%
gpt-oss
53.96
8.991.31
35.02
8.991.31
64.89%
100.00%
Llama 3.1 8B
23.58
3.932.47
17.78
3.932.47
75.40%
100.00%
Llama 3.1 70B
39.54
6.590.58
10.14
6.590.58
25.65%
100.00%
Mistral Small
38.87
6.481.66
17.58
6.481.66
45.23%
100.00%
Appendix
Table 7 : Statistics of generated intermediate questions (IQs) in ReHoPER on iLLC . We report the total number of generated IQs, the average number of IQs per path with standard deviation, the number of unique IQs, the average number of unique IQs per path with standard deviation, and the corresponding uniqueness percentages.
MuSR MM
Model
#IQ
#IQ/Path std
#uniqIQ
#uniqIQ/Path std
%uniqIQ
%uniqIQ/Path
Gemma 3
48.73
8.12 1.07
48.53
8.12 1.07
99.58%
100.00%
gpt-oss
59.70
9.95 0.10
59.62
9.95 0.10
99.87%
99.99%
Llama 3.1 8B
46.69
7.78 2.42
46.64
7.78 2.42
99.90%
99.97%
Llama 3.1 70B
50.04
8.34 1.20
49.74
8.34 1.20
99.41%
100.00%
Mistral Small
58.78
9.80 0.38
58.12
9.79 0.39
98.87%
99.94%
Appendix
Table 8 : Statistics of generated intermediate questions (IQs) in ReHoPER without replanning on MuSR MM, MoreHopQA, and iLLC L2-4. We report the total number of generated IQs, the average number of IQs per path with standard deviation, the number of unique IQs, the average number of unique IQs per path with standard deviation, and the corresponding uniqueness percentages.
MuSR MM
Model
#IQ
#IQ/Path std
#uniqIQ
#uniqIQ/Path std
%uniqIQ
%uniqIQ/Path
Gemma 3
60.00
10.00 0.00
58.76
9.95 0.09
97.93%
99.53%
gpt-oss
59.97
10.00 0.01
59.29
9.88 0.21
98.86%
98.88%
Llama 3.1 8B
60.00
10.00 0.00
57.86
9.69 0.47
96.43%
96.87%
Llama 3.1 70B
60.00
10.00 0.00
59.11
9.95 0.10
98.51%
99.49%
Mistral Small
60.00
10.00 0.00
57.53
9.62 0.52
95.88%
96.15%
Appendix
Table 9 : Statistics of generated intermediate questions (IQs) in ReHoPER without look-ahead on MuSR MM, MoreHopQA, and iLLC L2-4. We report the total number of generated IQs, the average number of IQs per path with standard deviation, the number of unique IQs, the average number of unique IQs per path with standard deviation, and the corresponding uniqueness percentages.
Gemma 3
MuSR MM (N: 250)
MoreHopQA (N: 150)
iLLC L2-4 (N: 250)
Step
T std
%T/N std
R std
%R/T std
%R/N std
T std
%T/N std
R std
%R/T std
%R/N std
T std
%T/N std
R std
%R/T std
%R/N std
1 → 2
250.00 0.00
100.0 0.0
249.83 0.37
99.9 0.1
99.9 0.1
138.17 2.27
92.1 1.5
122.50 3.69
88.7 3.0
81.7 2.5
250.00 0.00
100.0 0.0
209.33 21.02
83.7 8.4
83.7 8.4
2 → 3
250.00 0.00
100.0 0.0
250.00 0.00
100.0 0.0
100.0 0.0
124.33 2.49
82.9 1.7
105.50 2.14
84.9 2.6
70.3 1.4
250.00 0.00
100.0 0.0
131.00 34.24
52.4 13.7
52.4 13.7
3 → 4
250.00 0.00
100.0 0.0
249.83 0.37
99.9 0.1
99.9 0.1
118.17 3.13
78.8 2.1
101.33 3.68
85.8 2.0
67.6 2.5
248.17 1.21
99.3 0.5
73.83 16.87
29.7 6.8
29.5 6.7
4 → 5
250.00 0.00
100.0 0.0
250.00 0.00
100.0 0.0
100.0 0.0
106.00 5.32
70.7 3.5
90.50 3.91
85.4 2.3
60.3 2.6
241.17 7.60
96.5 3.0
74.83 18.96
31.1 8.0
29.9 7.6
Appendix
Table 10 : Trajectory statistics under the second-next analysis. T is the number of valid step transitions, R is the number of replanned transitions, and N is the total number of instances.
Llama 3.1 8B
MuSR MM (N: 250)
MoreHopQA (N: 150)
iLLC L2-4 (N: 250)
Step
T std
%T/N std
R std
%R/T std
%R/N std
T std
%T/N std
R std
%R/T std
%R/N std
T std
%T/N std
R std
%R/T std
%R/N std
1 → 2
250.00 0.00
100.0 0.0
250.00 0.00
100.0 0.0
100.0 0.0
109.67 5.12
73.1 3.4
101.83 4.52
92.9 1.4
67.9 3.0
221.83 8.63
88.7 3.5
219.17 9.10
98.8 1.0
87.7 3.6
2 → 3
250.00 0.00
100.0 0.0
249.00 1.00
99.6 0.4
99.6 0.4
105.67 4.15
70.4 2.8
99.00 6.14
93.6 2.8
66.0 4.1
210.33 9.62
84.1 3.8
199.00 9.27
94.6 2.3
79.6 3.7
3 → 4
249.67 0.47
99.9 0.2
248.83 0.90
99.7 0.3
99.5 0.4
94.67 2.49
63.1 1.7
88.67 2.21
93.7 2.3
59.1 1.5
201.83 11.91
80.7 4.8
186.50 9.39
92.5 1.0
74.6 3.8
4 → 5
250.00 0.00
100.0 0.0
249.00 0.82
99.6 0.3
99.6 0.3
79.67 2.29
53.1 1.5
74.17 2.91
93.1 1.7
49.4 1.9
176.50 13.46
70.6 5.4
154.50 13.51
87.5 1.9
61.8 5.4
Appendix
Table 11 : Trajectory statistics under the second-next analysis. T is the number of valid step transitions, R is the number of replanned transitions, and N is the total number of instances.
iLLC L2-4 (Development Set)
MoreHopQA
IO
CoT
PS
IO
CoT
PS
Model
T=.2
T=.3
T=.5
T=.7
T=.2
T=.3
T=.5
T=.7
T=.2
T=.3
T=.5
T=.7
T=.2
T=.3
T=.5
T=.7
T=.2
T=.3
T=.5
T=.7
T=.2
T=.3
T=.5
T=.7
Gemma 3
6.8
6.8
5.6
7.6
5.6
7.6
5.6
8.4
4.0
4.8
6.0
6.8
64.7
67.3
64.0
66.0
66.0
66.0
66.7
67.3
70.7
71.3
74.7
73.3
gpt-oss
90.8
92.4
93.2
94.8
89.2
92.0
93.6
94.8
89.6
92.8
91.2
93.6
78.7
78.7
81.3
80.0
78.0
79.3
80.0
80.0
80.0
80.7
82.0
79.3
Llama 3.1 8B
10.8
8.8
16.8
20.0
16.4
12.4
18.4
11.2
40.4
30.0
38.4
42.4
57.3
58.7
56.7
55.3
56.0
57.3
58.0
62.0
58.7
60.7
63.3
64.7
Mistral Small
8.0
7.2
9.2
7.2
5.6
8.8
9.2
6.0
8.8
8.4
8.4
6.4
66.0
70.7
71.3
70.0
67.3
68.0
67.3
72.0
70.7
70.7
70.7
68.0
Appendix
Table 12 : Accuracy (%) of ReHoPER under different sampling temperatures (T) on iLLC L2-4 (development set) and MoreHopQA, using IO, CoT, and PS answer-format prefixes. Green and light green mark the best and second best result, respectively, for each (model, dataset, prompting-style) triplet.
iLLC L2-4 (Development Set)
MoreHopQA
IO
CoT
PS
IO
CoT
PS
Model
5-12
6-10
10-6
12-5
5-12
6-10
10-6
12-5
5-12
6-10
10-6
12-5
5-12
6-10
10-6
12-5
5-12
6-10
10-6
12-5
5-12
6-10
10-6
12-5
Gemma 3
6.0
6.8
5.6
4.8
8.0
8.0
5.6
5.6
6.4
7.6
6.0
5.6
66.0
64.7
64.0
62.7
66.7
64.7
66.7
64.0
74.0
74.7
74.7
74.0
gpt-oss
94.0
94.4
93.2
94.8
93.6
93.6
93.6
93.6
93.2
94.4
91.2
92.8
78.0
78.0
81.3
80.0
78.0
81.3
80.0
79.3
79.3
82.0
82.0
78.7
Llama 3.1 8B
23.6
21.2
16.8
10.8
19.6
20.8
18.4
16.0
38.4
40.0
38.4
33.6
54.7
56.0
56.7
54.7
60.7
59.3
58.0
56.0
65.3
63.3
63.3
61.3
Mistral Small
10.0
10.0
9.2
5.6
12.8
12.8
9.2
8.0
11.2
9.6
8.4
6.4
70.7
71.3
71.3
72.0
67.3
67.3
67.3
67.3
71.3
70.7
70.7
71.3
Appendix
Table 13 : Accuracy (%) of ReHoPER under different { n }-{ mmax } configurations on iLLC L2-4 (development set) and MoreHopQA, using IO, CoT, and PS answer-format prefixes. Green and light green mark the best and second best result, respectively, for each (model, dataset, prompting-style) triplet.
MoreHopQA
iLLC L2-4
IO
CoT
PS
IO
CoT
PS
Model
SC
RH
SC
RH
SC
RH
SC
RH
SC
RH
SC
RH
Gemma 3
61.8 0.4
65.8 1.0
64.7 0.7
68.7 0.0
66.7 0.7
74.7 1.2
7.3 0.2
6.9 1.3
7.6 0.0
6.9 0.9
3.6 0.8
8.1 1.7
Llama 3.1 8B
60.7 1.2
57.1 1.7
60.0 0.7
60.0 1.8
62.4 1.5
63.1 0.4
11.9 0.8
17.5 4.1
12.5 3.2
18.4 2.0
30.4 1.1
41.6 1.4
Mistral Small
70.2 0.4
71.3 0.0
69.1 1.0
68.7 1.2
72.2 0.4
71.1 1.4
1.9 0.5
10.5 0.6
2.4 0.4
9.6 1.7
3.1 0.8
10.8 1.2
Phi-4
77.3 0.7
76.0 1.8
76.4 0.4
80.4 1.0
77.6 0.4
78.4 1.7
13.2 0.4
54.3 0.8
10.0 0.4
62.0 1.7
9.7 0.8
62.7 1.3
Appendix
Table 14 : Robustness of ReHoPER (RH) across repeated runs on MoreHopQA and iLLC L2-4 under IO, CoT, and PS answer-format prefixes. We report accuracy (%) as mean std over three runs. SC denotes simple question answering with Self-Consistency using the corresponding output prefix, while RH applies our planning approach but evaluates with the same prefix. Green , red , and gray indicate performance changes relative to SC.
iLLC L1-8
iLLC L2-8
IO
CoT
PS
IO
CoT
PS
Model
SD
SC
RH
SC
RH
SC
RH
SD
SC
RH
SC
RH
SC
RH
Gemma 3
26.4
7.6
12.4
36.0
28.8
52.8
32.8
0.0
0.4
0.0
0.0
0.0
0.0
0.0
gpt-oss
58.8
84.0
76.4
78.4
74.8
77.6
76.0
45.6
76.4
80.0
61.6
80.8
6.0
81.2
Llama 3.1 8B
7.6
31.2
27.6
24.0
39.2
39.2
34.4
0.0
0.8
2.0
0.0
4.0
0.4
2.4
Llama 3.1 70B
56.0
59.2
62.8
48.8
63.2
68.8
70.4
0.8
1.2
44.0
3.6
46.4
12.8
54.0
Appendix
Table 15 : Accuracy (%) of ReHoPER (RH) across iLLC L1-8 and iLLC L2-8 benchmarks under IO, CoT, and PS answer-format prefixes. SD refers to the Self-Discover baseline. Furthermore, SC denotes simple question answering with self-consistency using the corresponding output prefix, while RH applies our planning approach but evaluates with the same prefix. Green , red , and gray indicate performance changes relative to SC.
Table 16 : Models used in our experiments. We list the exact checkpoint, model family, parameter scale, and context length reported by the corresponding model card.
Figure 13 : Instruction Irh used for ReHoPER.
Figure 14 : Instruction Ifixed used for the ablation without replanning.
Figure 15 : Instruction Isingle used for the ablation without look-ahead.
Dataset
Task
Test Set Size
Answer Format
Per Task
MuSR
Murder Mysteries, Team Allocation
250
Choice number
Object Placement
256
MoreHopQA
-
150
Short text answer
iLLC
L1-4, L1-6, L1-8, L2-4, L2-6, L2-8
250
Concatenated letters
Appendix
Table 17: Summary of reasoning datasets used for evaluation.
While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Aggregation Layer dynamically reconciles these perspectives -- via final candidate selection, semantic synthesis, or neuro-symbolic verification -- to produce a robust global solution. We evaluate PoTRE on three frontier benchmarks: ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance. PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, surpassing the previous best official score. We demonstrate that this architectural heterogeneity achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines.
A common approach for teaching large language models (LLMs) to reason is to train on chain-of-thought (CoT) traces of in-distribution reasoning problems, but such annotated data is costly to obtain for every problem of interest. We want reasoning models to generalize beyond their training distribution, and ideally to generalize compositionally: combine atomic reasoning skills to solve harder, unseen reasoning tasks. We take a step towards compositional generalization of reasoning skills when addressing a target compositional task that has no labeled CoT data. We find that simply training models on CoT data of atomic tasks leads to limited generalization, but minimally modifying CoT formats of constituent atomic tasks to be composable can lead to improvements. We can train "atomic CoT" models on the atomic tasks with Composable CoT data and combine them with multitask learning or model merging for better zero-shot performance on the target compositional task. Such a combined model can be further bootstrapped on a small amount of compositional data using rejection sampling fine-tuning (RFT). Results on string operations and natural language skill compositions show that training LLMs on Composable CoT outperforms multitask learning and continued fine-tuning baselines within a given training data budget.
Fangcong Yin, Zeyu Leo Liu, Liu Leqi +2
The University of Texas at Austin · Princeton University
Multi-step reasoning remains a central challenge for large language models: single-pass generation is efficient but lacks accuracy; tree-search methods explore multiple paths but are computation-heavy. We address this gap by distilling reasoning progress into a hyperbolic geometric signal that guides step-by-step generation. Our approach is motivated by a structural observation: in combinatorial reasoning trees, solution-bearing states are few while dead ends are exponentially numerous. The hyperbolic space matches this asymmetry, with compact volume near the origin and exponentially expanding capacity toward the boundary, so that distance-to-origin naturally encodes solution proximity while angular separation distinguishes branches requiring different next operations. We train a lightweight head to project LLM hidden states into this space, then fine-tune a low-rank adapter interactively on its own reasoning attempts to act on the injected signal. Across multiple benchmarks, the geometric signal yields consistent gains, with larger improvements on deeper reasoning chains. Our code is publicly available at https://github.com/yuyuliu11037/HyperGuide.
Yuyu Liu, Haotian Xu, Yanan He +3
Department of Computer Science Stony Brook University · Department of Applied Mathematics and Statistics Stony Brook University · Department of Computer Science Yale University +2