Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how branching structures and sparse rewards induce these biases in long-horizon reasoning with local admissibility, and as a design principle for structural priors in less formal reasoning tasks. Motivated by this analysis, we propose SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning. SAGE combines two complementary structural guidance: algebraic sparsification, which projects locally admissible candidates onto operator-indexed algebraic subspaces to suppress spurious branching and mitigate exploration bias, and hyperbolic structural guidance, which embeds reasoning states into a negatively curved space to provide dense depth-wise signals and mitigate compounding bias. Across 12 benchmarks and 7 model families, SAGE outperforms competitive baselines. In particular, SAGE achieves up to an 8-fold improvement on the Andrews-Curtis problem, an open real-world long-horizon task. Code is available at: https://github.com/Susan571/SAGE-NeurIPS2026.
Figures & tables
Figure 1 : AC problem illustrates two long-horizon reasoning biases: exploration bias from many locally valid but low-promise branches, and compounding bias from early plausible deviations whose failures appear in later steps.
Figure 2 : Overview of SAGE workflow. Outcome-only training induces exploration and compounding biases; SAGE injects algebraic sparsification and hyperbolic structural guidance during RL, yielding focused search, stable reasoning, and no additional inference-time cost.
Model
MATH
Minerva
Olympiad
AIME24
AMC23
GSM8K
Putnam
Avg.
Flagship
Llama-3.3-70B-Instruct
66.08
33.61
32.94
17.31
29.02
80.47
9.91
38.48
2B models
Qwen3.5
50.64
11.62
24.58
9.41
43.36
47.38
3.66
27.24
Qwen3.5 w/SFT
60.41
26.52
27.96
3.74
37.88
55.06
4.63
30.89
Qwen3.5 w/GRPO
73.39
33.27
33.86
16.05
50.94
63.12
6.79
39.63
Table 1 : Accuracy (%) on mathematical reasoning benchmarks. Best in bold , second best underlined .
Model
STEM
MMLU-Pro
GPQA
BBH-H
ARC-C
Humanity
Social
Other
Avg.
9B models
Qwen3.5
12.71
8.02
14.95
10.21
11.06
10.91
21.58
18.49
Qwen3.5 w/SFT
20.18
11.36
28.97
19.22
19.85
12.31
32.44
29.56
Qwen3.5 w/GRPO
33.38
28.31
50.41
39.26
39.33
18.21
41.69
37.82
Qwen3.5 w/EMPO
32.57
27.32
48.73
37.69
37.91
21.11
44.04
39.73
Table 2 : Accuracy (%) on free-form natural reasoning benchmarks. The best is in bold with second best in underline .
Figure 3 : Comparative performance of models with SAGE versus base models across two primary metrics: AC Validity (Left) and Lean-Verified Proofs (Right). Numbers annotated above bars indicate the absolute percentage point improvement.
Variant
Olympiad
BBH-H
Acc
Rew./1k
Acc
Rew./1k
GRPO
38.82
16.08
41.69
20.06
EMPO
37.03
15.03
44.04
19.02
SAGE w/o ΨH
39.84
18.47
39.06
22.36
SAGE w/o ΨP
39.12
17.68
38.85
21.18
SAGE w/ Euclidean
38.01
16.92
38.02
20.89
Table 3: Core ablations on Qwen3.5-9B. All variants share the same entropy-filtered training subset, rollout budget, decoding constraints, optimization steps, and KL coefficient.
Method
Olym.
BBH-H
Lean
GRPO
65.31
63.29
14.64
GRPO-PRM
63.27
59.12
17.36
SAGE
68.26
69.07
23.69
Table 4 : Left: comparison with a learned process-reward baseline (Qwen3.5-35B). Right: component ablation on AC task (Qwen3.5-35B).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Hyperparameter ablation on the Olympiad dataset. Left: SAGE scales consistently across Pass@ K budgets. Middle: SAGE remains robust under different decoding temperatures. Right: SAGE is less sensitive to the KL coefficient βKL than GRPO.
Model
STEM
MMLU-Pro
GPQA
BBH-H
ARC-C
Humanity
Social
Other
Avg.
9B models
Qwen3.5
12.71 ± 0.21
8.02 ± 0.18
14.95 ± 0.25
10.21 ± 0.20
11.06 ± 0.16
10.91 ± 0.31
21.58 ± 0.38
18.49 ± 0.35
Qwen3.5 w/SFT
20.18 ± 0.34
11.36 ± 0.27
28.97 ± 0.43
19.22 ± 0.36
19.85 ± 0.31
12.31 ± 0.42
32.44 ± 0.55
29.56 ± 0.51
Qwen3.5 w/GRPO
33.38 ± 0.56
28.31 ± 0.49
50.41 ± 0.62
39.26 ± 0.55
39.33 ± 0.47
18.21 ± 0.73
41.69 ± 0.84
37.82 ± 0.79
Qwen3.5 w/EMPO
32.57 ± 0.53
27.32 ± 0.51
48.73 ± 0.66
37.69 ± 0.58
37.91 ± 0.50
21.11 ± 0.69
44.04 ± 0.82
39.73 ± 0.76
Appendix
Table 5 : Accuracy (%) on free-form natural reasoning benchmarks over 5 random seeds, reported as mean with standard deviation . The best mean is in bold with second best in underline .