Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how branching structures and sparse rewards induce these biases in long-horizon reasoning with local admissibility, and as a design principle for structural priors in less formal reasoning tasks. Motivated by this analysis, we propose SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning. SAGE combines two complementary structural guidance: algebraic sparsification, which projects locally admissible candidates onto operator-indexed algebraic subspaces to suppress spurious branching and mitigate exploration bias, and hyperbolic structural guidance, which embeds reasoning states into a negatively curved space to provide dense depth-wise signals and mitigate compounding bias. Across 12 benchmarks and 7 model families, SAGE outperforms competitive baselines. In particular, SAGE achieves up to an 8-fold improvement on the Andrews-Curtis problem, an open real-world long-horizon task. Code is available at: https://github.com/Susan571/SAGE-NeurIPS2026.
Figures & tables
Figure 1 : AC problem illustrates two long-horizon reasoning biases: exploration bias from many locally valid but low-promise branches, and compounding bias from early plausible deviations whose failures appear in later steps.
Figure 2 : Overview of SAGE workflow. Outcome-only training induces exploration and compounding biases; SAGE injects algebraic sparsification and hyperbolic structural guidance during RL, yielding focused search, stable reasoning, and no additional inference-time cost.
Model
MATH
Minerva
Olympiad
AIME24
AMC23
GSM8K
Putnam
Avg.
Flagship
Llama-3.3-70B-Instruct
66.08
33.61
32.94
17.31
29.02
80.47
9.91
38.48
2B models
Qwen3.5
50.64
11.62
24.58
9.41
43.36
47.38
3.66
27.24
Qwen3.5 w/SFT
60.41
26.52
27.96
3.74
37.88
55.06
4.63
30.89
Qwen3.5 w/GRPO
73.39
33.27
33.86
16.05
50.94
63.12
6.79
39.63
Table 1 : Accuracy (%) on mathematical reasoning benchmarks. Best in bold , second best underlined .
Model
STEM
MMLU-Pro
GPQA
BBH-H
ARC-C
Humanity
Social
Other
Avg.
9B models
Qwen3.5
12.71
8.02
14.95
10.21
11.06
10.91
21.58
18.49
Qwen3.5 w/SFT
20.18
11.36
28.97
19.22
19.85
12.31
32.44
29.56
Qwen3.5 w/GRPO
33.38
28.31
50.41
39.26
39.33
18.21
41.69
37.82
Qwen3.5 w/EMPO
32.57
27.32
48.73
37.69
37.91
21.11
44.04
39.73
Table 2 : Accuracy (%) on free-form natural reasoning benchmarks. The best is in bold with second best in underline .
Figure 3 : Comparative performance of models with SAGE versus base models across two primary metrics: AC Validity (Left) and Lean-Verified Proofs (Right). Numbers annotated above bars indicate the absolute percentage point improvement.
Variant
Olympiad
BBH-H
Acc
Rew./1k
Acc
Rew./1k
GRPO
38.82
16.08
41.69
20.06
EMPO
37.03
15.03
44.04
19.02
SAGE w/o ΨH
39.84
18.47
39.06
22.36
SAGE w/o ΨP
39.12
17.68
38.85
21.18
SAGE w/ Euclidean
38.01
16.92
38.02
20.89
Table 3: Core ablations on Qwen3.5-9B. All variants share the same entropy-filtered training subset, rollout budget, decoding constraints, optimization steps, and KL coefficient.
Method
Olym.
BBH-H
Lean
GRPO
65.31
63.29
14.64
GRPO-PRM
63.27
59.12
17.36
SAGE
68.26
69.07
23.69
Table 4 : Left: comparison with a learned process-reward baseline (Qwen3.5-35B). Right: component ablation on AC task (Qwen3.5-35B).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Hyperparameter ablation on the Olympiad dataset. Left: SAGE scales consistently across Pass@ K budgets. Middle: SAGE remains robust under different decoding temperatures. Right: SAGE is less sensitive to the KL coefficient βKL than GRPO.
Model
STEM
MMLU-Pro
GPQA
BBH-H
ARC-C
Humanity
Social
Other
Avg.
9B models
Qwen3.5
12.71 ± 0.21
8.02 ± 0.18
14.95 ± 0.25
10.21 ± 0.20
11.06 ± 0.16
10.91 ± 0.31
21.58 ± 0.38
18.49 ± 0.35
Qwen3.5 w/SFT
20.18 ± 0.34
11.36 ± 0.27
28.97 ± 0.43
19.22 ± 0.36
19.85 ± 0.31
12.31 ± 0.42
32.44 ± 0.55
29.56 ± 0.51
Qwen3.5 w/GRPO
33.38 ± 0.56
28.31 ± 0.49
50.41 ± 0.62
39.26 ± 0.55
39.33 ± 0.47
18.21 ± 0.73
41.69 ± 0.84
37.82 ± 0.79
Qwen3.5 w/EMPO
32.57 ± 0.53
27.32 ± 0.51
48.73 ± 0.66
37.69 ± 0.58
37.91 ± 0.50
21.11 ± 0.69
44.04 ± 0.82
39.73 ± 0.76
Appendix
Table 5 : Accuracy (%) on free-form natural reasoning benchmarks over 5 random seeds, reported as mean with standard deviation . The best mean is in bold with second best in underline .
Large language models (LLMs) achieve strong reasoning performance by allocating substantial computation at inference time, often generating long and verbose reasoning traces. While recent work on efficient reasoning reduces this overhead through length-based rewards or pruning, many approaches are post-trained under a much shorter context window than base-model training, a factor whose effect has not been systematically isolated. We first show that short-context post-training alone, using standard GRPO without any length-aware objective, already induces substantial reasoning compression-but at the cost of increasingly unstable training dynamics and accuracy degradation. To address this, we propose Step-level Advantage Selection (SAS), which operates at the reasoning-step level and assigns a zero advantage to low-confidence steps in correct rollouts and to high-confidence steps in verifier-failed rollouts, where failures often arise from truncation or verifier issues rather than incorrect reasoning. Across diverse mathematical and general reasoning benchmarks, SAS improves average Pass@1 accuracy by 0.86 points over the strongest length-aware baseline while reducing average reasoning length by 16.3%, yielding a better accuracy-efficiency trade-off.
Recent studies observe that reinforcement learning with verifiable rewards (RLVR) reliably improves pass@1 on reasoning tasks, yet often fails to yield comparable gains in pass@k, raising the question of whether RLVR genuinely enables large language models to acquire novel reasoning abilities or merely enhances the efficiency of sampling reasoning modes already present in the base model. Prior analyses largely support the latter view, attributing this limitation to structural properties of standard RLVR objectives that result in insufficient exploration pressure. In this work, we argue that a central structural constraint arises from reverse-KL regularization, which stabilizes training but inherently anchors the policy to the reference distribution, thereby suppressing the emergence of alternative reasoning modes. However, we show that neither removing the KL term nor replacing it with forward-KL provides a satisfactory solution, as both disrupt the efficiency-coverage trade-off by either inducing reward hacking or allocating probability mass to off-target regions. To resolve this tension, we propose SAGE, a principled framework that enables controllable empirical support expansion by reshaping the reverse-KL anchor distribution itself through a guide function q(x,y), achieving consistent improvements in both pass@1 and pass@k across challenging mathematical reasoning benchmarks. Our code is available at https://github.com/tally0818/SAGE.
Long-horizon agentic reasoning requires large language models to act over long interaction histories containing thoughts, tool calls, observations, and partial conclusions. The challenge is not merely that these histories grow long, but that information needed for the current decision may be scattered across distant steps and only become relevant later. Existing approaches address this difficulty by truncating the interaction history, compressing it into shorter surrogates, or retrieving selected parts of it for reuse, but they do not explicitly model how access to past interaction should adapt to the agent's evolving state. We instead cast long-horizon reasoning as a problem of state-adaptive memory. To this end, we propose State-Adaptive Memory~(SAM), a standalone framework that consolidates ongoing interaction into compact memory cues while preserving raw trajectory pages for intent-driven recall. These cues are not treated as replacements for history; rather, they serve as lightweight handles that allow the agent to reconstruct temporally distant information according to its current needs, without retraining the underlying backbone. We further optimize the memory module through expert-guided supervision and reinforcement learning, aligning it with trajectory-level utility. Across BrowseComp, BrowseComp-ZH, WideSearch, and HLE, SAM consistently outperforms strong baselines over diverse agent backbones. Our results suggest that explicit memory modeling provides a simple and effective foundation for long-horizon agentic reasoning.
Yuyang Hu, Hongjin Qian, Shuting Wang +5
GSAI, Renmin University of China · Beijing Academy of Artificial Intelligence