When a model solves a problem on one attempt and fails it on the next, what separates the two is rarely the final answer token; it is the trajectory that reached it. Contrastive activation steering leaves that signal unused: CAA, SADI, RepE and ITI build their direction from experimenter-supplied text, recorded while the model reads rather than reasons. That choice also caps what the vector can express, since polarity must be written into the text, and a task judged only by outcome offers nothing to write it with. ROAST makes the trajectory itself the contrast: sample rollouts, let an outcome verifier split them into successes and failures, and contrast the reasoning that worked against the reasoning that did not. A matched teacher-forced control---rollouts, labels, answer text and pair counts held fixed, the trajectory alone stripped---points to the trajectory as what matters: on GSM8K at 0.6B the pairs alone buy +0.12 points while restoring the trajectories buys +6.05, the larger and only seed-robust step. Replacing the trajectory with an equal-length neutral prefix or another question's reasoning falls below no intervention. The two corpora are also far apart geometrically, a median 70+ degrees apart at both Qwen3 scales probed, beyond what a split-half null explains. Reading from rollouts calls for two corrections---keeping the full difference vector rather than Top-10% masking, and giving each question one vote rather than one per pair---and only grouped aggregation beats the unsteered baseline under 20% verifier noise. On parser-free benchmarks (GSM8K, MATH500, IFEval), ROAST is best in all six cells over two models, by up to +9.7, at +6.4% wall-clock and no added context; it also leads on six parser-scored benchmarks across three models. Across nine models (0.6B--122B, four families), ROAST improves on the unsteered model at every scale. Code: https://github.com/TomySu404/ORBIT
Figures & tables
Model
IFEval
TruthfulQA
Model
IFEval
TruthfulQA
Qwen3-0.6B
62.46 ± 1.27
48.78 ± 0.52
Qwen3-32B
80.31 ± 0.14
85.39 ± 0.21
+ ROAST
63.27 ± 0.28
49.91 ± 0.23
+ ROAST
81.02 ± 0.71
86.64 ± 0.21
Qwen3-8B
77.35 ± 0.32
78.66 ± 0.16
GLM4-32B
80.31 ± 0.71
57.40 ± 1.66
+ ROAST
80.50 ± 1.98
81.62 ± 0.03
+ ROAST
81.11 ± 0.71
69.51 ± 3.83
GLM4-9B
81.02 ± 0.12
53.83 ± 0.01
Llama-3-70B
81.02 ± 0.44
85.20 ± 0.61
+ ROAST
81.82 ± 0.42
60.10 ± 0.17
+ ROAST
82.19 ± 0.53
86.83 ± 0.70
Table 1: Scaling across capacity and architecture, including a 70 B dense model and a 122 B mixture-of-experts model. Accuracy (%) ± std; bold marks the steered result. The sign is consistent—all 16 cells improve—but not the size : these two tasks leave little headroom (six of eight unsteered IFEval scores already in 79.9 – 84.4 ) and demand the least reasoning of our nine benchmarks, the two conditions under which § 5.3 predicts small gains; the large margins are on reasoning (Table 4 ). All cells use the standard configuration of § 5.1 (grouped, L2 , N=1000 ; IFEval N=100 throughout).
Method
Polarity source
Trajectory in pair
Native outcome pair
CAA
gold vs. wrong option
no (answer only)
no
RepE
honest vs. dishonest instr.
no (stimulus shared)
no †
ITI
curated true/false statements
no (statement only)
no
SADI
appended label
no (answer only)
no
ROAST
outcome verifier
yes (full rollout)
yes
Table 2: What carries the contrast, and what follows, as each method is published —not a claim that no adaptation exists (§ 5.3 tests one). Native outcome pair asks whether the method’s own construction yields a pair, without an appended label, on a task judged only by outcome (GSM8K, MATH500, MMLU, IFEval). † RepE’s construction applies syntactically—an instruction pair needs no gold answer—but no instruction expresses “solve this correctly” versus “solve this incorrectly,” so the pair is not polarized by the outcome; we pair opposing instruction framings instead (App. B.7 ).
Figure 1: The extraction–deployment mismatch. (a) What each method reads h from: prior methods, a curated ground-truth answer; ROAST, the model’s own rollout, which contains the reasoning. The two rows differ in text , not in how the forward pass is run (§ 3.1 ). (b) Qwen3-0.6B/GSM8K, 64 questions, 8 rollouts each: cos(μtf,μar) has median 0.28 over early/middle layers and relative L2 exceeds 1 in 92% of all layers; the same probe on Qwen3-8B gives a median angle of 72∘ . Panel (b) compares two readings of one target direction , not a good-vs-bad behavior gap, and it ranks neither.
Figure 2: Geometric view of Grouped Mean Normalization. Step 1: per-question normalization collapses each question to a unit direction vˉq , discarding prompt-specific scale. Step 2a: their Euclidean mean vavg lies inside the ball, its norm measuring directional agreement. Step 2b: the outer normalization projects it back to the sphere, so α alone controls strength.
Model
Method
SST2
SST5
MMLU
TruthfulQA
Winogrande
XNLI
Avg
Qwen3-0.6B
No intervention
78.80 ± 0.29
25.65 ± 0.30
35.49 ± 0.01
48.78 ± 0.52
49.80 ± 0.59
37.54 ± 0.12
46.01
Few-shot (100)
84.03 ± 0.38
30.07 ± 1.63
40.20 ± 0.05
48.21 ± 0.11
49.49 ± 0.08
44.15 ± 0.02
49.36
CAA
78.55 ± 0.25
28.73 ± 0.02
35.81 ± 0.01
35.19 ± 0.00
49.80 ± 0.00
37.13 ± 0.01
44.20
SADI
82.81 ± 0.03
28.08 ± 0.03
38.16 ± 0.04
24.13 ± 0.08
49.85 ± 0.06
43.92 ± 0.02
44.49
RepE
82.37 ± 0.55
28.51 ± 0.61
37.60 ± 0.28
47.05 ± 0.40
50.44 ± 0.39
43.16 ± 0.55
48.19
ROAST
84.60 ± 0.38
27.91 ± 0.14
38.15 ± 0.01
49.91 ± 0.23
52.36 ± 0.00
41.72 ± 0.01
49.11
Table 3: Main results. Accuracy (%) ± std over three seeds. Bold / underline : best and second-best steering method per column within a model; few-shot ICL in italics is a reference only, consuming context at inference. All methods use one fixed configuration and N=1000 (ROAST: grouped, L2 ); N=100 and nongrouped results are in Appendix C.7 .
Qwen3-0.6B
Qwen3-8B
Method
GSM8K
MATH500
IFEval
GSM8K
MATH500
IFEval
No intervention
48.86 ± 0.19
33.59 ± 0.47
62.46 ± 1.27
92.61 ± 0.66
56.09 ± 0.78
77.35 ± 0.32
Few-shot (20)
46.21 ± 0.17
34.19 ± 0.16
–
92.46 ± 0.75
57.38 ± 0.35
–
CAA
48.56 ± 0.66
34.12 ± 1.42
–
90.06 ± 0.09
52.56 ± 1.22
–
SADI
56.30 ± 0.85
36.25 ± 2.34
–
91.95 ± 0.38
55.94 ± 0.94
–
RepE
53.44 ± 0.63
35.18 ± 1.05
62.98 ± 0.49
91.63 ± 0.27
56.10 ± 0.79
78.66 ± 1.22
Table 4: Reasoning and instruction following. Accuracy (%) ± std over three seeds; bold is the best steering method, underline second. Few-shot ( italics ) is capped at 20 exemplars by context length and does not apply to IFEval, whose prompts are themselves the constraint.
Qwen3-0.6B
Qwen3-8B
#
Variant
SST2
MMLU
GSM8K
MATH500
SST2
MMLU
GSM8K
MATH500
1
Teacher-forced + Top- 10%
79.57 ± 0.17
37.78 ± 0.19
48.74 ± 0.10
34.51 ± 0.28
86.24 ± 0.20
67.38 ± 0.05
89.40 ± 0.65
53.34 ± 0.26
2
+ matched ROC pairs (still TF)
–
37.82 ± 0.17
48.86 ± 0.54
34.81 ± 0.32
–
67.51 ± 0.17
89.94 ± 0.19
53.64 ± 0.24
3
ROC + Top- 10%
81.95 ± 0.15
39.41 ± 0.32
54.91 ± 0.46
36.02 ± 0.36
88.18 ± 0.48
68.22 ± 0.17
91.80 ± 0.21
56.17 ± 0.05
4
ROC + full + nongrouped
82.29 ± 0.37
39.57 ± 0.19
55.39 ± 0.85
36.13 ± 0.12
87.75 ± 0.26
68.68 ± 0.11
91.67 ± 0.23
57.41 ± 0.15
5
+ grouped mean
83.08 ± 0.12
40.71 ± 0.17
57.17 ± 0.51
36.30 ± 0.21
88.16 ± 0.20
68.92 ± 0.16
92.42 ± 0.50
57.23 ± 0.54
Table 5: Component isolation with a matched teacher-forced control. Each row toggles one axis relative to the row above: extraction (TF → ROC), scaling (Top- k→ full), aggregation (nongrouped → grouped, mean → max →L2 ); rows 5–7 each replace row 4’s nongrouped aggregation with one grouped reducer. Row 2 matches ROC on rollouts, labels, answer text, and pair count, stripping only the trajectory. Accuracy (%) ± std over the same three seeds ( 33,42,52 ) as the main tables. All rows hold N=100 and come from one re-run of the ablation grid, so cells are comparable within this table but not with Table 3 or with the other single-seed N=100 tables (App. C.13 ); conclusions here are internal to N=100 (App. C.7 ).
Extraction context
Accuracy
vs. no intervention
No intervention
48.60
—
Answer only (matched TF)
43.47
−5.13
C1
+ length-matched filler
32.87
−15.73
C2
+ mismatched trajectory
36.67
−11.93
C3
+ shuffled own trajectory
46.27
−2.33
+ intact own trajectory (full ROC)
57.60
+9.00
Table 6: Extraction-context controls on GSM8K (Qwen3-0.6B, N=100 , three seeds, grouped L2 , all arms at α=3 ). The five steered arms carry the same token count and a byte-identical anchor token, and differ only in how much of this question’s reasoning the context retains. Configuration internal to this table (App. C.13 ).
GSM8K
TruthfulQA
Flips
grp
nongrp
grp
nongrp
0%
58.52
57.63
49.22
47.47
5%
58.01
55.90
48.74
45.31
10%
57.06
52.88
47.95
42.10
20%
54.83
46.71
46.02
37.88
Table 7: Left: verifier noise. Symmetric label flips in the R+/R− partition (Qwen3-0.6B, N=100 , single seed; grp = grouped, red = below the unsteered baseline ( 48.86 GSM8K, 48.78 TruthfulQA); App. C.5 ). Right: runtime (min) , Qwen3-8B/MMLU on one H20, at each method’s own grid ( 4 points for ROAST, 6 CAA, 24 SADI; App. C.11 ).
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: ROAST end to end. Step 1 (ROC, § 4.1 ): the model produces n rollouts per question under par ; a verifier marks each correct (green) or incorrect (red), and the activation at the final token of each rollout is retained at every layer. Step 2 (CSS + grouped aggregation, § 4.2 – 4.3 ): the difference of means within a question is normalized to a unit direction, those directions are averaged over questions, and the result is normalized once more. Step 3 (intervention, § 4.4 ): at inference the vector, scaled by α , is added to component c of every layer l while a new input is decoded. Only Step 3 is online. (The figure labels Step 1 “on-distribution”; we call it on-policy in the text, following § E on the scope of that term.)
Dataset
Total eval
Steering set
Dev / Test
Source split
SST2
872
100/1000
174 / 698
dev
SST5
2,210
100/1000
442 / 1,768
test
MMLU
11,173
100/1000
2,234 / 8,939
test
TruthfulQA
717
100/1000
143 / 574
validation
Winogrande
1,267
100/1000
253 / 1,014
dev
XNLI
5,010
100/1000
1,002 / 4,008
test
Appendix
Table 8: Dataset splits. Total eval excludes the steering set. Steering-set size lists the N=100 ablation and the N=1000 main setting.
Model
Task
gen
gen margin
lp
lp margin
fail (un. → ROAST)
Qwen3-0.6B
MMLU
35.10→42.30
+7.20
36.10→35.80
−0.30
2.6→0.7
TruthfulQA
35.01→34.59
−0.42
34.73→34.73
0.00
0.0→0.0
Winogrande
49.10→49.00
−0.10
48.90→48.70
−0.20
0.0→0.3
Qwen3-8B
MMLU
72.20→68.40
−3.80
74.90→75.00
+0.10
3.6→2.8
TruthfulQA
70.29→70.29
0.00
66.25→66.11
−0.14
0.0→0.0
Winogrande
65.20→51.30
−13.90
64.90→64.50
−0.40
0.5→23.8
Appendix
Table 9: Preliminary log-probability rescoring (seed 42 , α=3 fixed rather than dev-selected, N=1000 at 0.6 B and N=300 at 8 B). gen is the generate-then-extract protocol; lp is log-probability ranking, which has no parser; fail is the extraction-failure rate. This configuration is not that of Table 3 —its gen margins do not reproduce that table and on 8 B/MMLU differ in sign—so the table below should be read as a diagnostic of this run, not as a re-measurement of our reported margins.
Qwen3-0.6B
Qwen3-8B
Comparison
What it measures
cos
angle
cos
angle
Random directions
chance floor ( d−1/2 )
0.00±.031
90.0∘
0.00±.016
90.0∘
Within-TF (half A/B)
reliability of the TF reading
0.9957
5.3∘
0.9938
6.4∘
Within-AR (half A/B)
reliability of the AR reading
0.9720
13.6∘
0.9434
19.4∘
Cross TF vs. AR
the reported mismatch
0.2738 [ .18,.47 ]
74.1∘
0.3044 [ .27,.67 ]
72.3∘
disattenuated
mismatch net of noise
0.2775
73.9∘
0.3145
71.7∘
Appendix
Table 10: Split-half null. GSM8K, 64 questions ×8 rollouts, median over the first two thirds of layers (IQR in brackets). Within-TF and within-AR re-estimate one corpus’s direction on two disjoint halves of the questions; cross is the comparison of § 3.1 . The disattenuated row divides the cross cosine by ρtfρar .
Figure 4: Per-layer view of Table 10 . Two estimates of one corpus (blue, green) track each other near 1 at every layer, while the cross-corpus reading (red) runs far below them and far above the chance floor (dotted). Both readings converge again in the final layers, which is why medians are reported over the first two thirds throughout.
Model
Task
All correct
All incorrect
Discarded
Usable
Qwen3-0.6B
SST2
33.5%
3.0%
36.5%
63.5%
Qwen3-0.6B
MATH500
14.0%
40.0%
54.0%
46.0%
Qwen3-8B
SST2
79.5%
6.5%
86.0%
14.0%
Qwen3-8B
MATH500
49.5%
26.0%
75.5%
24.5%
Appendix
Table 11: Fraction of steering questions with no valid contrastive pair ( G=8 rollouts, N=200 prompts per cell, temperature 1.0 , top- p0.95 ). This is a standalone diagnostic run at a higher-entropy sampling setting than the main configuration ( T=0.8 , top- p0.9 ; § 5.1 ), which makes rollouts more likely to split, so the discard rates here are if anything optimistic relative to the main runs. It is a diagnostic of the regime, not a measurement of the realized ∣Q∣ in any table.
Steering vector
SST2
GSM8K
MMLU
None (unsteered)
78.80 / 86.68
48.86 / 92.61
35.49 / 66.99
SST2 vector
85.24 / 89.94
48.71 / 92.55
35.42 / 66.90
GSM8K vector
78.60 / 86.51
58.52 / 92.95
35.55 / 67.10
MMLU vector
78.74 / 86.60
48.60 / 92.48
40.73 / 69.11
Appendix
Table 12: Cross-task interference: apply task A ’s vector (rows) while evaluating task B (columns), as 0.6 B / 8 B accuracy (%). Each vector improves its own task (diagonal) while off-diagonal changes remain within the scale of the seed-to-seed variation reported in the main-paper tables. Single seed; the baseline row is repeated from the main-paper tables.
Model
SST2
MMLU
TruthfulQA
Grouped avg
Nongrouped avg
Qwen3-0.6B, N=100
85.24 / 85.03
40.73 / 37.32
49.22 / 47.47
50.03
49.22
Qwen3-0.6B, N=1000
84.60 / 84.81
38.15 / 37.32
49.91 / 49.39
49.11
48.52
Qwen3-8B, N=100
89.94 / 89.11
67.71 / 69.06
81.18 / 80.14
68.79
69.48
Qwen3-8B, N=1000
88.47 / 89.11
69.11 / 69.31
81.62 / 79.70
70.22
69.33
Gemma3-1B, N=100
81.95 / 81.23
39.87 / 39.68
19.43 / 19.86
47.26
46.86
Gemma3-1B, N=1000
84.31 / 83.10
40.20 / 39.91
23.47 / 22.77
47.82
48.09
Appendix
Table 13: Grouped versus nongrouped aggregation at two steering-set sizes. Each cell is grouped / nongrouped accuracy (%). The main paper reports the grouped, L2 , N=1000 configuration throughout; N=100 rows are included only to isolate steering-set size.
n
SST2
MATH500
2
80.72
36.25
8
84.66
36.41
64
85.24
36.31
128
85.74
36.78
Appendix
Table 14: Rollout budget and anchor position (Qwen3-0.6B), both at the ablation setting N=100 and so not comparable cell-for-cell with the main tables (App. C.13 ). Left: rollouts per question. Right: extraction anchor, token indices counted from the first response token; dashes were not run. Entries with ± are over three seeds, the rest single seed.
Component
GSM8K
MATH500
IFEval
Layer scope
SST2
MATH500
MLP
59.86
36.41
63.27
All
85.24
36.41
Attention
48.34
34.69
57.93
First 5
81.38
32.34
MLP + Attn.
49.67
36.88
60.34
Last 5
81.02
33.13
Appendix
Table 15: Intervention placement (Qwen3-0.6B, single seed). Left: which component receives the vector. Right: which layers it is applied to. MLP at all layers is best.
Figure 5: Rollout stability on Qwen3-0.6B. The extracted direction converges as the rollout budget grows, while per-layer L2 norms remain stable across rollout budgets.
Task
Unsteered
α=0.5
α=1.0
α=3.0
α=5.0
Qwen3-0.6B
MMLU
36.12
37.74
38.32
39.12
42.79
SST2
77.59
82.18
85.06
87.93
86.78
SST5
24.21
29.86
31.67
31.00
28.73
TruthfulQA
50.35
52.45
52.45
52.45
53.15
Winogrande
46.64
53.36
53.75
53.36
54.15
Appendix
Table 16: Dev-set accuracy (%) across steering strengths α , single seed. Moderate values are best; too-large α can collapse behavior (Qwen3-8B MMLU at α=5 ).
Figure 6: Per-layer L2 norm of the aggregated steering vector (log scale). Ungrouped aggregation inflates with depth ( 0.87→72.4 and 0.93→27.0 ) so magnitude is governed by depth, not signal; grouping stays flat, leaving α in control.
Figure 7: Per-question magnitude ∥Δhq∥ and directional consistency, each normalized by its mean (mid-network layer). Magnitude varies substantially across questions yet is only weakly and negatively correlated with consistency ( r=−0.10 to −0.44 ), so a magnitude-weighted average adds variance without adding signal.
Figure 8: Discrete Top- k masking discards about half the signal. Left: cumulative L2 energy of the contrastive difference vector as a function of the fraction of coordinates retained; keeping the top 10% retains only ∼48% of the energy. Right: the same quantity per layer, which stays near 50% throughout the network on both models.
Figure 9: Task and layer specificity. Left: inter-layer cosine similarity of steering vectors (MMLU); distinct layers are nearly orthogonal (scale ±0.06 ), so a single global direction would be inadequate and layer-wise intervention is necessary. Right: cross-dataset cosine similarity (Qwen3-8B); even closely related tasks such as SST2 and SST5 share little direction ( <0.1 ), indicating distribution-specific rather than category-level features.
Figure 10: Per-layer norm profiles of steering vectors for MLP versus attention components, which have distinct magnitude structure across depth.