Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: https://github.com/caiyuchen-ustc/On_Policy_Vector_Training
Figures & tables
Figure 1: The single steering vector reveals that RLVR-induced changes form a low-dimensional, depth-sensitive control geometry, characterized by (a) effective manifold capacity and (b) control-manifold separation. Building on these findings, (c) Alpha-Stabler stabilizes training via principal-subspace intrusion monitoring and backward gradient projection.
Figure 2: (a) On-policy distillation scores on the training set. (b) Off-policy distillation scores on the evaluation set. (c) Scores of vector steering under varying teacher types. (d) Scores of full-parameter training, LoRA, and vector steering under off-policy distillation at three distribution gaps.
Figure 3: Layer-wise performance. (a) Converged performance of steering vectors across injection layers. (b) Task performance during training for steering vectors. (c) Steering gradient norms at different layers. (d) KL divergence with the teacher under fixed versus gated steering. (e) Converged performance of fixed versus gated injection.
Figure 4: Diagnosing and mitigating high-layer degradation. (a) Converged performance of vector steering under different nonlinear parameters. (b) Layer-wise mean pairwise cosine similarity for hidden states and output gradients. (c) Distribution of gating activations at different layers.
Figure 5: Multi-vector steering with sequential orthogonalization. (a) Training trajectories of sequentially learned orthogonal steering vectors on SciKnowEval with DeepSeek-R1-Distill-Qwen-1.5B. (b) Standalone accuracy of the extracted directions across model scales. Each direction is trained and evaluated independently, with preceding directions used only to define the orthogonality constraint.
Figure 6: Geometric consistency and transfer of steering directions. (a) Layer-averaged cosine similarities between steering vectors under different teacher parameterizations, shown as stacked bars for each reference vector va(i) , with segments ordered from v(0) to v(3) indicating similarities with {vb(j)}j=03 . (b) Layer-averaged cosine similarities between on- and off-policy steering vectors with matching extraction indices, per task. (c) Cross-topic directional alignment versus normalized source-to-target transfer gains, with a fitted regression line.
Figure 7: Steering geometry and principal-subspace intrusion. Energy fractions in (a) and (b) use the Top- 10% principal subspace. (a) Standalone accuracy and principal-subspace energy fractions of sequentially extracted directions across four scientific topics. (b) The same quantities when steering vectors are restricted to the Top- 30% activation-PC subspace. (c) Task score and PSI during RL training, comparing stable and collapse-prone runs.
Figure 8: (a, b) Training dynamics with and without Alpha-Stabler. (c) Training dynamics across learning rates. (d) Task score and peak PSI under different intervention timings. (e) Performance under control of different subspaces.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Training Set
Evaluation Set
Reward
Math
DeepMath-103K ( He et al., 2025 )
AIME 2024 ( Ye et al., 2025 )
Answer matching
MATH500 ( Lightman et al., 2023 )
Science
SciKnowEval-Training ( Feng et al., 2024 )
SciKnowEval-Eval ( Feng et al., 2024 )
Answer matching
Code
Eurus ( Cui et al., 2025a )
LiveCodeBench-v5 ( Jain et al., 2024 )
Unit tests
Instruction following
IFEval-Training ( Zhao et al., 2024 )
IFEval-Eval ( Zhao et al., 2024 )
Constraint checking
Appendix
Table 1: Training and evaluation datasets and reward functions for each domain.
RL
OPD
Off-PD
LR, full fine-tuning
1×10−6
1×10−5
5×10−5
LR, LoRA ( r=8 )
3×10−5
5×10−6
1×10−5
LR, vector steering
—
1×10−1
5×10−2
Max prompt / response
2,048 / 20,480
3,072 / 16,384
3,072 / 16,384
Prompt batch
128
1,024
1,024
Mini-batch (per step)
32
1,024
1,024
Appendix
Table 2: Learning rates and paradigm-specific hyperparameters.
Figure 9: Vocabulary-space readout of steering vectors across network depth. (a) Normalized entropy of softmax(RMSNorm(vL)WU⊤) . The dashed line at 0.9498 denotes the mean entropy of normalized Gaussian directions passed through the same readout pipeline. (b) Highest-logit vocabulary items for steering vectors learned at representative layers. All vectors are normalized before readout.
Figure 10: Effect of single-vector steering on hidden-state geometry in Qwen3-4B ( k=20 , H=2560 ), measured at layer L+1 for injection at layer L . (a) Mean principal-subspace overlap between clean and steered activations. (b) Ratio of the top- k explained variance after and before steering. (c) Relative displacement of the activation centroid. Dashed lines denote the corresponding reference values.
Figure 11: Multi-vector steering across all four task domains and model scales. The first vector is consistently the most effective, with subsequent directions showing gradually decreasing standalone performance; the decay is slower for larger models.
Figure 12: Comparison of pre-collapse score trajectories with and without Alpha-Stabler at learning rates of (a) 10−4 , (b) 5×10−5 , and (c) 10−5 . Shaded regions denote the first 50 warm-up updates of Alpha-Stabler. Controlled and uncontrolled runs exhibit essentially the same score growth before collapse, indicating that Alpha-Stabler’s stabilization introduces almost no additional computational overhead.
Const.
Name
Meaning
Related analysis
ε
principal intrusion
Upper bound on target principal-to-complement energy in the sensitivity metric
PC-energy measurements under compatible metrics and target definitions
η
complement residual
Residual-to-shared energy within the complement
Context dependence and fixed-vector approximation
κ
amplitude consistency
1/(1+\CV(a)2)
Consistency of the shared target amplitude
Appendix
Table 3: Geometric quantities entering the sufficient recovery bound. Their empirical counterparts require matching shift definitions, distributions, and metrics.
Symbol
Type
Meaning
x∼D
context
Token position and prefix, x=(q,y<t)
q , Y
sequence level
Prompt and complete response, including termination
hℓ(x)∈Rd
ctx.-dep.
Layer- ℓ hidden state
z0(x),zT(x)∈RV
ctx.-dep.
Base and teacher logits
t(x)=zT(x)−z0(x)
ctx.-dep.
Teacher logit shift modulo 1V ; Remark D.3
Fx∈RV×V
ctx.-dep.
Fisher information at teacher logits; Equation ( 38 )
Appendix
Table 4: Recurring symbols for the distillation analysis. Context dependence is indicated by (x) or a subscript x ; Part III separately introduces batched training quantities.
Statement
Reference
Scope / status
Part I — fixed-vector recovery
Distillation admits a weighted least-squares approximation
Proposition D.1
Fixed contexts; controlled network and KL remainders
The regularized optimum is an exponential reward tilt
Lemma D.2
Exact optimum of the stated objective, β>0
Next-token ratios follow continuation normalizers
Lemma D.3
Exact under the tilt model
Shared profiles and consistent pullbacks give shared shifts
Lemmas D.4 , D.5
Additional sufficient structural conditions
ρ=SNR/(1+SNR)
Theorem D.1
Constant-metric surrogate target energy
Appendix
Table 5: Summary of recovery, steering geometry, and gradient-control results, with their assumptions and empirical scope.
Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models (LLMs), yet it remains limited by sparse binary rewards and its ignorance of model-internal uncertainty. In this paper, we propose ConSteer-RL, a simple yet effective framework that integrates token-level confidence signals derived from model log-probabilities into RLVR training. Specifically, building upon the Group Relative Policy Optimization (GRPO) framework, we construct a confidence-aware reward by aggregating per-token probabilities into a scalar confidence score and incorporating it into an awareness-based reward shaping mechanism that penalizes overconfident errors while reinforcing correct and confident reasoning. Experimental results demonstrate that ConSteer-RL consistently outperforms strong GRPO baselines, achieving average improvements of 2.3%-4.0% across different model scales.
Qing Miao, Yiming Zhao, Jing Yang +5
Xi’an Jiaotong University · University of Science and Technology of China
Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the underlying geometry of the resulting parameter trajectories remains underexplored. In this work, we demonstrate that RLVR weight trajectories are extremely low-rank and highly predictable. Specifically, we find that the majority of downstream performance gains are captured by a rank-1 approximation of the parameter deltas, where the magnitude of this projection evolves near-linearly with training steps. Motivated by this, we propose a simple and compute-efficient method RELEX (REinforcement Learning EXtrapolation), which estimates the rank-1 subspace from a short observation window and extrapolates future checkpoints via linear regression, with no learned model required. Across three models (i.e., Qwen2.5-Math-1.5B, Qwen3-4B-Base, and Qwen3-8B-Base), RELEX produces checkpoints that match or exceed RLVR performance on both in-domain and out-of-domain benchmarks, requiring as few as 15% steps of full RLVR training. Remarkably, RELEX is able to extrapolate far beyond the observation window at no training cost, predicting checkpoints up to 10-20× beyond the observed prefix with continued improvement (e.g., observe only the first 50 steps and extrapolate to 1000 steps). Our ablation analysis confirms the minimalist sufficiency of RELEX: neither increasing the subspace rank nor employing non-linear modeling yields further gains in extrapolation. Finally, we show that RELEX's success stems from a "denoising" effect: by projecting updates onto the rank-1 subspace, the model discards stochastic optimization noise that would otherwise degrade performance during extrapolation. Our code is available at https://github.com/weizhepei/RELEX.
Zhepei Wei, Xinyu Zhu, Wei-Lin Chen +3
University of Virginia · Washington University in St. Louis
Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR intrinsically relies on ground-truth labels for reward computation, the acquisition of which is often prohibitively expensive in real-world scenarios. While unsupervised RLVR paradigms attempt to circumvent this by training on pseudo-labels, they are notoriously susceptible to training collapse. Moreover, different samples often exhibit varying annotation values. In this paper, we propose Reinforcement Learning with Active Verifiable Rewards (RLAVR), which actively acquires ground-truth labels for a small set of selected samples and integrates them with pseudo-labels, thereby stabilizing training dynamics and improving performance under limited annotation budgets. To identify valuable samples, we propose the Corrective Advantage Gap (CAG) metric and analyze the sample-level supervision value. Building on this, we introduce Correction-Aware Reliability Estimation for RLAVR (CARE), which translates the oracle CAG criterion into a practical pre-query acquisition policy to substantially improve training stability. Extensive experiments across diverse domains, model families, and model scales demonstrate the effectiveness and generality of our approach. Our code is available at https://github.com/Lumina04/CARE.
Li Wang, Xiaodong Lu, Xiaohan Wang +5
Meituan · Beihang University · Nanyang Technological University