Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contrast or learning signal. Prior hint-based methods construct auxiliary hints from solution evidence and use them to re-solve failed queries, recovering learning signal. Yet the resulting trajectories are typically treated as ordinary solution trajectories despite being generated under an assisted condition unavailable at evaluation. We discover hinted reward shift: recovered reward contrast can concentrate policy updates on hinted trajectories, limiting improvement without hints. This also creates a trade-off: increasing hinted trajectories can accelerate early learning but intensify reward shift later. To address this problem, we propose HATCH (Hint-Annealed Self-Teaching), an online single-policy framework that learns from both generating and using its own hints to improve reasoning without assistance. To mitigate hinted reward shift, we introduce online weighting to anneal the contribution of hinted trajectories. However, learning to generate hints can conflict with improving query solving. We therefore use gradient projection to remove the opposing component of hint-generation updates. Together, these designs support self-improvement by enabling the policy to create learning opportunities for itself and turn them into stronger reasoning without hints. We evaluate our method on mathematical reasoning benchmarks and outperform state-of-the-art methods by 1.02 pp on Llama-3.2-1B-Instruct, 2.84 pp on Qwen3-1.7B, and 4.32 pp on Qwen3-8B.
Figures & tables
Figure 1: Comparison between GRPO, hint-based GRPO, and ours. (a) GRPO receives no group-relative learning signal for hard queries with all-incorrect rollouts. (b) Hint-based GRPO restores learning signal by re-solving under a hint, yielding strong improvement during training but limited improvement during testing. (c) Ours learns from generating and using its own hints, while gradually shifting training toward the original solving ability without hints.
Figure 2: Evidence of hinted reward shift on Qwen3-8B. (a) Left: Mean absolute advantages of original and hinted trajectories. (b) Middle: Gradient-angle distribution between original and hinted solving over steps 1–50. (c) Right: Original-query accuracy with different numbers of hints.
Figure 3: Overview of our method. For solve-none queries, it first self-generates hints from solutions and then re-solves the query with the hints, yielding three complementary training trajectories: native solving, hint generation, and hinted solving. Online weighting adaptively controls the contribution of hinted-solving feedback, while gradient projection removes conflicting components from hint-generation updates. The resulting signals are jointly used to update a single shared policy.
Method
Math500
Minerva
Oly.
AIME24/25/26
AMC23
HMMT25
BRUMO25
Avg.
Llama-3.2-1B-Instruct
24.04
4.41
4.71
1.11/0.00/0.00
5.00
0.00
1.11
4.49
+ GRPO
25.45
5.02
4.71
2.22 / 1.11 / 1.11
6.67
1.11
1.11
5.39
+ QuESTA
26.92
6.13
5.21
2.22 / 1.11 / 1.11
7.50
2.22
2.22
6.07
+ HiLL
26.12
5.63
4.96
1.11 / 2.22 / 1.11
8.33
0.00
2.22
5.74
+ HATCH (ours)
27.38
6.62
5.65
2.22 / 2.22 / 3.33
10.83
2.22
3.33
7.09
Qwen3-1.7B
79.76
52.57
48.66
16.67/16.67/10.00
42.50
6.67
16.67
32.24
Table 1: Per-dataset no-hint accuracy (%) for the initial models and each trained method. Bold marks the best result and underlining marks the second-best distinct result with ties retained.
Figure 5: Left: average benchmark accuracy over training updates. Middle: best-so-far no-hint benchmark accuracy against cumulative GPU-hours for the two online methods. Right: accuracy of training by our method with single-, two-, and four retained hinted-solve groups.
Variant
Acc. (%)
HATCH (ours)
53.76
Offline Hints
48.07
External Hints
48.80
Online Hints
50.16
Fixed Weighting
44.96
Linear Decay
48.25
Table 2: Ablation results on Qwen3-8B.
Figure 6: Ablation dynamics on Qwen3-8B. Left: relative strength in the solution direction in the final shared update gradient. Middle: auxiliary-weight schedules induced by the three γ settings. Right: the corresponding nine-benchmark accuracy curves of each selected γ .
Figure 7: No-hint group-state dynamics across Ours, HiLL, and GRPO on Qwen3-8B. Ours progressively converts solve-none groups into solve-partial and solve-all groups, indicating stronger unassisted solving. From left to right panels are: solve-none, solve-partial and solve-all groups.
Reinforcement Learning (RL) has become a key driver for enhancing the long chain-of-thought (CoT) reasoning capabilities of Large Language Models (LLMs). However, prevalent methods like GRPO often fail when task difficulty exceeds model capacity, leading to reward sparsity and inefficient training. Prior work attempts to mitigate this with off-policy data, but such methods often induce severe distributional mismatches that destabilize policy updates. In this work, we identify a core issue underlying these failures, which we term low training affinity, and introduce Affinity, the first quantitative metric for monitoring the compatibility between external guidance and the model's intrinsic policy. To address this, we propose HINT, an adaptive framework designed to enhance reasoning capabilities while explicitly preserving high Affinity. First, instead of revealing partial answers, HINT supplies Meta-Hints, which act as abstract cognitive scaffolding to guide the model in articulating solutions independently. Second, to ensure stability, we integrate Affinity-Aware Policy Optimization (AAPO), which dynamically modulates the learning objective based on the Affinity. Extensive experiments across diverse benchmarks demonstrate that HINT consistently outperforms strong baselines, while exhibiting superior stability and robust generalization to out-of-distribution tasks. Code is available at https://github.com/ViviqwerAsd/HINT.
Xinyi Wang, Jinyi Han, Zishang Jiang +7
School of Data Science, Fudan University · 2Shanghai Institute of Artificial Intelligence for Education, East China Normal University · College of Computer Science and Artificial Intelligence, Fudan University +1
Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization direction but rarely explain how a model should revise its mistaken reasoning, which can encourage shortcut learning and brittle policies. We propose \textbf{SocraticPO} (Socratic Policy Optimization), a policy-optimization framework that augments RL rollouts with Socratic-style natural-language guidance. During rollout, the student first answers independently; if the answer is incorrect, a teacher diagnoses the attempt and provides concise corrective guidance, after which the student continues under the expanded context. Crucially, this guidance is paired with reward decay: correct answers obtained after teacher intervention only receive decayed rewards, preventing the policy from treating teacher help as a free path to reward. Since SocraticPO only modifies the rollout process while leaving the standard expected-reward objective intact, it can be plugged into existing policy-gradient backends such as Reinforce++. Moreover, because the teacher provides only text-level guidance, SocraticPO can leverage stronger black-box teacher models without requiring access to logits or distribution matching. On undergraduate-level scientific reasoning benchmarks from SciKnowEval, SocraticPO improves over strong RL and self-distillation baselines. Ablations show that both targeted guidance and reward decay are necessary, with reward decay mitigating reliance on assisted correction.
Zirui Liu, Jie Ouyang, Qi Liu +8
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China · iFLYTEK AI Research (Central China), iFLYTEK Co., Ltd
Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Rewards (RLVR) being a promising enhancement strategy. However, existing reward mechanisms are constrained to the outcome-level correctness and lack explicit signals to guide the model to consider diverse solutions. In contrast, human problem solving typically involves evaluating multiple potential approaches and selecting the most reliable solution, a cognitive process that current RLVR frameworks do not explicitly incentivize. Inspired by this, we propose Hint-Guided Diversified Policy Optimization (HDPO), allowing the model to first list all potential candidate solution outlines as hints and then select the most reliable one for further reasoning. HDPO comprises two stages of Cold Start for Structured Reasoning and Hint-Guided Diversified Reinforcement Learning to incentivize the model to generate diverse and reliable solutions following the ``propose-select-think'' trajectory. Experimental results show that HDPO effectively boosts LLM reasoning and enhances the diversity of candidate solutions as well as the LLM's ability to identify reliable solutions.
Zhiyu Cao, Kaixin Wu, Mingjie Zhong +4
School of Computer Science and Technology, Soochow University, Suzhou, China · Ant Group, Hangzhou, China