Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: https://github.com/caiyuchen-ustc/On_Policy_Vector_Training
Figures & tables
Figure 1: The single steering vector reveals that RLVR-induced changes form a low-dimensional, depth-sensitive control geometry, characterized by (a) effective manifold capacity and (b) control-manifold separation. Building on these findings, (c) Alpha-Stabler stabilizes training via principal-subspace intrusion monitoring and backward gradient projection.
Figure 2: (a) On-policy distillation scores on the training set. (b) Off-policy distillation scores on the evaluation set. (c) Scores of vector steering under varying teacher types. (d) Scores of full-parameter training, LoRA, and vector steering under off-policy distillation at three distribution gaps.
Figure 3: Layer-wise performance. (a) Converged performance of steering vectors across injection layers. (b) Task performance during training for steering vectors. (c) Steering gradient norms at different layers. (d) KL divergence with the teacher under fixed versus gated steering. (e) Converged performance of fixed versus gated injection.
Figure 4: Diagnosing and mitigating high-layer degradation. (a) Converged performance of vector steering under different nonlinear parameters. (b) Layer-wise mean pairwise cosine similarity for hidden states and output gradients. (c) Distribution of gating activations at different layers.
Figure 5: Multi-vector steering with sequential orthogonalization. (a) Training trajectories of sequentially learned orthogonal steering vectors on SciKnowEval with DeepSeek-R1-Distill-Qwen-1.5B. (b) Standalone accuracy of the extracted directions across model scales. Each direction is trained and evaluated independently, with preceding directions used only to define the orthogonality constraint.
Figure 6: Geometric consistency and transfer of steering directions. (a) Layer-averaged cosine similarities between steering vectors under different teacher parameterizations, shown as stacked bars for each reference vector va(i) , with segments ordered from v(0) to v(3) indicating similarities with {vb(j)}j=03 . (b) Layer-averaged cosine similarities between on- and off-policy steering vectors with matching extraction indices, per task. (c) Cross-topic directional alignment versus normalized source-to-target transfer gains, with a fitted regression line.
Figure 7: Steering geometry and principal-subspace intrusion. Energy fractions in (a) and (b) use the Top- 10% principal subspace. (a) Standalone accuracy and principal-subspace energy fractions of sequentially extracted directions across four scientific topics. (b) The same quantities when steering vectors are restricted to the Top- 30% activation-PC subspace. (c) Task score and PSI during RL training, comparing stable and collapse-prone runs.
Figure 8: (a, b) Training dynamics with and without Alpha-Stabler. (c) Training dynamics across learning rates. (d) Task score and peak PSI under different intervention timings. (e) Performance under control of different subspaces.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Training Set
Evaluation Set
Reward
Math
DeepMath-103K ( He et al., 2025 )
AIME 2024 ( Ye et al., 2025 )
Answer matching
MATH500 ( Lightman et al., 2023 )
Science
SciKnowEval-Training ( Feng et al., 2024 )
SciKnowEval-Eval ( Feng et al., 2024 )
Answer matching
Code
Eurus ( Cui et al., 2025a )
LiveCodeBench-v5 ( Jain et al., 2024 )
Unit tests
Instruction following
IFEval-Training ( Zhao et al., 2024 )
IFEval-Eval ( Zhao et al., 2024 )
Constraint checking
Appendix
Table 1: Training and evaluation datasets and reward functions for each domain.
RL
OPD
Off-PD
LR, full fine-tuning
1×10−6
1×10−5
5×10−5
LR, LoRA ( r=8 )
3×10−5
5×10−6
1×10−5
LR, vector steering
—
1×10−1
5×10−2
Max prompt / response
2,048 / 20,480
3,072 / 16,384
3,072 / 16,384
Prompt batch
128
1,024
1,024
Mini-batch (per step)
32
1,024
1,024
Appendix
Table 2: Learning rates and paradigm-specific hyperparameters.
Figure 9: Vocabulary-space readout of steering vectors across network depth. (a) Normalized entropy of softmax(RMSNorm(vL)WU⊤) . The dashed line at 0.9498 denotes the mean entropy of normalized Gaussian directions passed through the same readout pipeline. (b) Highest-logit vocabulary items for steering vectors learned at representative layers. All vectors are normalized before readout.
Figure 10: Effect of single-vector steering on hidden-state geometry in Qwen3-4B ( k=20 , H=2560 ), measured at layer L+1 for injection at layer L . (a) Mean principal-subspace overlap between clean and steered activations. (b) Ratio of the top- k explained variance after and before steering. (c) Relative displacement of the activation centroid. Dashed lines denote the corresponding reference values.
Figure 11: Multi-vector steering across all four task domains and model scales. The first vector is consistently the most effective, with subsequent directions showing gradually decreasing standalone performance; the decay is slower for larger models.
Figure 12: Comparison of pre-collapse score trajectories with and without Alpha-Stabler at learning rates of (a) 10−4 , (b) 5×10−5 , and (c) 10−5 . Shaded regions denote the first 50 warm-up updates of Alpha-Stabler. Controlled and uncontrolled runs exhibit essentially the same score growth before collapse, indicating that Alpha-Stabler’s stabilization introduces almost no additional computational overhead.
Const.
Name
Meaning
Related analysis
ε
principal intrusion
Upper bound on target principal-to-complement energy in the sensitivity metric
PC-energy measurements under compatible metrics and target definitions
η
complement residual
Residual-to-shared energy within the complement
Context dependence and fixed-vector approximation
κ
amplitude consistency
1/(1+\CV(a)2)
Consistency of the shared target amplitude
Appendix
Table 3: Geometric quantities entering the sufficient recovery bound. Their empirical counterparts require matching shift definitions, distributions, and metrics.
Symbol
Type
Meaning
x∼D
context
Token position and prefix, x=(q,y<t)
q , Y
sequence level
Prompt and complete response, including termination
hℓ(x)∈Rd
ctx.-dep.
Layer- ℓ hidden state
z0(x),zT(x)∈RV
ctx.-dep.
Base and teacher logits
t(x)=zT(x)−z0(x)
ctx.-dep.
Teacher logit shift modulo 1V ; Remark D.3
Fx∈RV×V
ctx.-dep.
Fisher information at teacher logits; Equation ( 38 )
Appendix
Table 4: Recurring symbols for the distillation analysis. Context dependence is indicated by (x) or a subscript x ; Part III separately introduces batched training quantities.
Statement
Reference
Scope / status
Part I — fixed-vector recovery
Distillation admits a weighted least-squares approximation
Proposition D.1
Fixed contexts; controlled network and KL remainders
The regularized optimum is an exponential reward tilt
Lemma D.2
Exact optimum of the stated objective, β>0
Next-token ratios follow continuation normalizers
Lemma D.3
Exact under the tilt model
Shared profiles and consistent pullbacks give shared shifts
Lemmas D.4 , D.5
Additional sufficient structural conditions
ρ=SNR/(1+SNR)
Theorem D.1
Constant-metric surrogate target energy
Appendix
Table 5: Summary of recovery, steering geometry, and gradient-control results, with their assumptions and empirical scope.