cs.LGFeb 15, 2026

Watch the Model Think: On-Policy Extraction of Activation Steering Vectors

Authors: Xuanbo Su, Yingfang Zhang, Hao Luo, Huajun Bai, Guangyuan Dong, Ling Huang

Organizations: Bairong Inc., Beijing, China · School of Mathematics, Harbin Institute of Technology, Harbin, China

Abstract

When a model solves a problem on one attempt and fails it on the next, what separates the two is rarely the final answer token; it is the trajectory that reached it. Contrastive activation steering leaves that signal unused: CAA, SADI, RepE and ITI build their direction from experimenter-supplied text, recorded while the model reads rather than reasons. That choice also caps what the vector can express, since polarity must be written into the text, and a task judged only by outcome offers nothing to write it with. ROAST makes the trajectory itself the contrast: sample rollouts, let an outcome verifier split them into successes and failures, and contrast the reasoning that worked against the reasoning that did not. A matched teacher-forced control---rollouts, labels, answer text and pair counts held fixed, the trajectory alone stripped---points to the trajectory as what matters: on GSM8K at 0.6B the pairs alone buy +0.12 points while restoring the trajectories buys +6.05, the larger and only seed-robust step. Replacing the trajectory with an equal-length neutral prefix or another question's reasoning falls below no intervention. The two corpora are also far apart geometrically, a median 70+ degrees apart at both Qwen3 scales probed, beyond what a split-half null explains. Reading from rollouts calls for two corrections---keeping the full difference vector rather than Top-10% masking, and giving each question one vote rather than one per pair---and only grouped aggregation beats the unsteered baseline under 20% verifier noise. On parser-free benchmarks (GSM8K, MATH500, IFEval), ROAST is best in all six cells over two models, by up to +9.7, at +6.4% wall-clock and no added context; it also leads on six parser-scored benchmarks across three models. Across nine models (0.6B--122B, four families), ROAST improves on the unsteered model at every scale. Code: https://github.com/TomySu404/ORBIT

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Predicting Future Behaviors in Reasoning Models Enables Better Steering

    Jun 9, 2026Evgenii Kortukov, Piotr Komorowski, Florian Klein +5Large Reasoning ModelsLinear Activation Steering

  2. What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

    Apr 9, 2026Stephen Cheng, Sarah Wiegreffe, Dinesh ManochaLinear Activation SteeringSteering