Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.
Figures & tables
Safety outcomes (lower is better)
AIM ASR ↓
Refusal Suppression ASR ↓
Over-refusal ↓
Method
Pre
Post
Δ
Pre
Post
Δ
Pre
Post
Δ
Base
0.411
0.735
+0.324
0.288
0.394
+0.106
0.205
0.169
-0.036
DPO
0.027
0.436
+0.409
0.017
0.149
+0.132
0.277
0.175
-0.103
CAA
0.407
0.739
+0.332
0.291
0.398
+0.107
0.217
0.164
-0.053
Probe
0.409
0.735
+0.326
0.289
0.391
+0.102
0.203
0.164
-0.039
Table 1: Safety control before and after benign fine-tuning. Pre and Post denote results before and after fine-tuning on Alpaca-Cleaned, and Δ=Post−Pre . Lower ASR and over-refusal (OR) are better. Results are unweighted macro-averages across three base models. Direct-prompting ASR is omitted because it is generally lower than under jailbreak attacks; per-model values are reported in Appendix Table 12 . Full per-model results are provided in Appendix B.1 and B.2 .
MMLU ↑
GSM8K ↑
HumanEval ↑
Method
Pre
Post
Δ
Pre
Post
Δ
Pre
Post
Δ
Base
0.659
0.661
+0.002
0.788
0.703
-0.085
0.747
0.800
+0.053
DPO
0.653
0.666
+0.013
0.790
0.687
-0.103
0.713
0.800
+0.087
CAA
0.663
0.659
-0.003
0.790
0.715
-0.075
0.753
0.793
+0.040
Probe
0.657
0.661
+0.004
0.788
0.707
-0.082
0.760
0.807
+0.047
Flow
0.616
0.600
-0.016
0.782
0.778
-0.003
0.793
0.787
-0.007
Table 2: Capability scores before and after benign fine-tuning. Pre and Post denote scores before and after fine-tuning on Alpaca-Cleaned, and Δ=Post−Pre . Scores are unweighted macro-averages across three base models; higher scores and positive changes are better. Full per-model results are provided in Appendix B.2 .
Figure 1: Data efficiency and cross-scope safety transfer. (a) Refusal-suppression ASR across training-data sizes and sources on Qwen2.5-1.5B-Instruct; lower is better. Error bars show the standard deviation over three seeds for 100- and 200-pair settings; larger settings are single runs. (b) Each cell reports the unweighted macro-average ASR change across three base models, averaging three independent seeds within each model. Rows denote evaluation scopes and columns denote intervention-training scopes. Outlined cells mark matched-scope evaluation; negative changes indicate improved safety relative to the base model.
Full-response Detection
Streaming Early Detection
Monitor
AUROC
AUPRC
Cal.-1%
Cal.-5%
Recall
Median Pos.
Seq. FPR
Representation probes
Mean
0.982
0.978
0.896/0.039
0.986/0.190
0.908
0.291
0.081
Last
0.972
0.966
0.737/0.025
0.954/0.104
0.618
0.257
0.145
Rolling
0.971
0.966
0.928/0.068
0.962/0.122
0.913
0.143
0.017
Attention
0.965
0.932
0.855/0.039
0.968/0.145
0.824
0.154
0.047
Table 3: Monitoring performance on native Qwen2.5-32B-Instruct trajectories. For full-response detection, Cal.- x% reports test TPR/realized FPR at a threshold selected to target x% FPR on the calibration set. For streaming detection, Recall is the fraction of harmful responses detected, Median Pos. is the normalized first-detection position among detected responses, and Seq. FPR is the fraction of safe trajectories that trigger at least one alarm. Higher AUROC, AUPRC, TPR, and Recall are better; lower FPR and Med. Pos. are better. Bold denotes the best overall result, and underline denotes the best representation-probe result.
Figure 2: Native monitoring efficiency and replay-based control integration. Left: Probes that reuse native Qwen2.5-32B-Instruct activations provide competitive detection at substantially lower marginal cost than text monitors. Right: After benign fine-tuning, probe-triggered blocking reduces the mean jailbreak ASR of matched-data DPO on Qwen2.5-14B-Instruct from 0.500 to 0.042, approaching its pre-fine-tuning ASR of 0.058.
Figure 3: Three monitor–control integration strategies. (a) Blocking replaces a flagged response with a refusal. (b) Regeneration uses the monitor’s alarm to request a safe rewrite of a flagged draft. (c) Stronger steering increases the flow-intervention strength and reruns generation.
Figure 4: Monitor–control integration before and after benign fine-tuning. Curves show mean ASR across AIM and refusal suppression and over-refusal on Qwen2.5-14B-Instruct as the alarm threshold varies. Horizontal dashed lines denote the corresponding DPO and Flow baselines without monitoring; vertical dotted lines mark the rolling-probe thresholds previously calibrated at 1% and 5% FPR.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Pairs
Train/Val.
Epochs
Learning rate
Batch
DPO
5,406
5,406/–
2
5×10−5
128–256
CAA
5,406
–
–
–
–
Probe
5,406
4,866/540
200
10−2
full
Flow
5,406
5,406/–
10
3×10−4
4–8
Appendix
Table 4: Training configuration of the control methods. Batch size is effective batch size for DPO and per-device batch size for Flow.
Base model
DPO batch
CAA/Probe layer
α
Flow layer
T
Qwen2.5-1.5B
8×32
24
1.0
20
0.7
Llama-3.1-8B
8×16
28
1.0
16
0.7
Qwen2.5-14B
4×32
31
1.0
24
0.7
Appendix
Table 5: Model-specific control settings. DPO batch is reported as per-device batch × gradient-accumulation steps. CAA and Probe share the same intervention layer and strength.
Scope
Pair pool
Pairs used
Eval. prompts
Cybercrime
787
787
67
Physical harm
1,571
1,500
75
Toxicity
2,316
1,500
75
Appendix
Table 6: Domain-specific data used in the granularity evaluation. Pool is the number of complete safe–unsafe pairs available; all methods use the same number of pairs in each scope.
Rubric
Accuracy
FP
FPR
FN
FNR
Legacy StrongREJECT
0.905
29
0.112
30
0.082
Binary-v2
0.960
24
0.093
1
0.003
Binary-v3
0.986
6
0.023
3
0.008
Appendix
Table 7: Validation of the harmful-compliance judge against the 624-example author-adjudicated reference set. FPR is computed over safe responses and FNR over harmful responses.
Model
Method
Layer
Base
α=0.5
α=1
α=2
α=4
Qwen2.5-1.5B
CAA
8
0.606/0.220
0.590/0.220
0.588/0.215
0.600/0.221
0.580/0.221
12
0.592/0.210
0.608/0.205
0.574/0.180
0.514/0.165
16
0.600/0.210
0.596/0.210
0.582/0.215
0.536/0.200
20
0.600/0.215
0.598/0.215
0.596/0.195
0.586/0.190
24
0.610/0.210
0.598/0.200
0.594/0.195
0.608/0.220
PROBE
8
0.606/0.215
0.598/0.205
0.602/0.210
0.614/0.195
0.626/0.195
Appendix
Table 8: Complete layer–strength sweep for CAA and probe-based steering. Each entry reports ASR/OR, where lower is better for both metrics. Bold entries denote the operating points used in the main experiments.
Model
Layer
Base
T=0.5
T=0.7
T=1
T=1.5
Qwen2.5-1.5B
8
0.488/0.240
0.476/0.280
0.482/0.275
0.452/0.280
0.450/0.285
12
0.484/0.235
0.428/0.270
0.388/0.275
0.320/0.265
0.262/0.290
16
0.498/0.240
0.422/0.295
0.396/0.295
0.329/0.295
0.308/0.300
20
0.492/0.240
0.466/0.240
0.454/0.245
0.428/0.270
0.384/0.270
24
0.486/0.230
0.486/0.260
0.494/0.260
0.464/0.265
0.418/0.285
Llama-3.1-8B
8
0.396/0.175
0.306/0.165
0.266/0.120
0.260/0.110
0.474/0.100
Appendix
Table 9: Layer–strength sweep for flow-based steering. Each entry reports ASR/OR, where lower is better for both metrics. We report the four central intervention strengths and omit the two endpoint stress-test settings ( T=0.3 and T=2 ) for compactness. Bold entries denote the operating points used in the main experiments.
AIM ASR
Refusal Suppression ASR
Over-refusal
Layer
Pre
Post
Δ
Pre
Post
Δ
Pre
Post
Δ
16 (Primary)
0.153
0.099
-0.054
0.144
0.259
+0.115
0.148
0.172
+0.024
8 (Post hoc)
0.077
0.712
+0.636
0.224
0.466
+0.243
0.124
0.156
+0.032
Appendix
Table 10: Flow layer sensitivity to benign fine-tuning on Llama-3.1-8B-Instruct. Both layers use the full 5,406-pair, 10-epoch Flow training configuration with T=0.7 . Pre and Post denote results before and after fine-tuning on Alpaca-Cleaned, and Δ=Post−Pre . Lower ASR and over-refusal (OR) are better. Layer 16 is the operating point used in the main comparison; Layer 8 is evaluated post hoc following the layer sweep.
AIM
Refusal Suppression
Model
Method
Pre
Post
Δ
Pre
Post
Δ
Qwen2.5-1.5B
Base
0.833
0.849
+0.016
0.349
0.519
+0.170
DPO
0.013
0.410
+0.397
0.003
0.170
+0.167
CAA
0.818
0.853
+0.035
0.349
0.511
+0.162
Probe
0.827
0.849
+0.022
0.349
0.513
+0.164
Flow
0.652
0.780
+0.128
0.256
0.463
+0.207
Appendix
Table 11: Per-model robustness before and after benign fine-tuning. Pre and Post denote ASR before and after fine-tuning on Alpaca-Cleaned, and Δ=Post−Pre . Lower is better.
Model
Method
Pre
Post
Qwen2.5-1.5B
Base
0.061
0.204
DPO
0.000
0.080
CAA
0.071
0.195
Probe
0.067
0.208
Flow
0.042
0.125
Llama-3.1-8B
Base
0.019
0.010
Appendix
Table 12: Direct-prompting harmful-compliance ASR. Pre and Post denote results before and after fine-tuning on Alpaca-Cleaned. Each entry is computed from 312 or 313 valid binary-v3 judgments; Macro is the unweighted mean across the three base models. Lower is better.
Over-refusal
MMLU
GSM8K
HumanEval
Model
Method
Pre
Post
Δ
Pre
Post
Δ
Pre
Post
Δ
Pre
Post
Δ
Qwen2.5-1.5B
Base
0.476
0.304
-0.172
0.576
0.572
-0.004
0.615
0.410
-0.205
0.660
0.660
0.000
DPO
0.692
0.324
-0.368
0.582
0.576
-0.006
0.625
0.365
-0.260
0.660
0.700
+0.040
CAA
0.500
0.288
-0.212
0.582
0.568
-0.014
0.635
0.455
-0.180
0.660
0.660
0.000
Probe
0.460
0.300
-0.160
0.578
0.572
-0.006
0.630
0.420
-0.210
0.660
0.680
+0.020
Flow
0.632
0.636
+0.004
0.432
0.432
0.000
0.615
0.615
0.000
0.620
0.620
0.000
Appendix
Table 13: Per-model practicality before and after benign fine-tuning. Lower over-refusal (OR) and higher capability scores are better. Δ=Post−Pre .
Data source
Pairs
DPO
CAA
Probe
Flow
Full data
100
0.346±0.006
0.343±0.010
0.346±0.012
0.364±0.004
200
0.347±0.005
0.344±0.002
0.346±0.008
0.354±0.005
500
0.333
0.340
0.343
0.337
1k
0.314
0.346
0.346
0.330
2k
0.263
0.343
0.340
0.301
5k
0.250
0.353
0.349
0.353
Appendix
Table 14: Data-efficiency results on Qwen2.5-1.5B-Instruct. Entries are refusal-suppression ASR (lower is better). Values for 100 and 200 pairs are mean ± standard deviation over three seeds; larger settings are single runs.
Method
Training data
Pairs
AIM
RS
OR
DPO
Original
63,094
0.029
0.125
0.092
DPO
Contrastive
5,406
0.000
0.000
0.088
Flow
Contrastive
5,406
0.153
0.144
0.148
Appendix
Table 15: Data-source comparison on Llama-3.1-8B-Instruct. The main DPO and Flow configurations use the same 5,406 filtered contrastive pairs; DPO trained on the complete original preference set is included as a reference. We report ASR under AIM and refusal suppression (RS), together with over-refusal (OR). Lower is better.
Model
Method
Evaluation scope
Base
General
Cybercrime
Physical harm
Toxicity
Qwen2.5-1.5B
DPO
General
0.281
0.173
0.218
0.178
0.169
Cybercrime
0.582
0.363
0.463
0.343
0.338
Physical harm
0.453
0.289
0.400
0.302
0.320
Toxicity
0.200
0.111
0.160
0.133
0.120
CAA
General
0.281
0.254
0.254
0.254
0.254
Cybercrime
0.582
0.567
0.582
0.582
0.582
Appendix
Table 16: Per-model granularity and cross-scope transfer. Entries are mean ASR over three independent runs with different random seeds under refusal suppression (lower is better). Rows denote evaluation scopes and columns denote intervention-training scopes; bold values are matched-scope diagonal entries. Base ASRs are listed once for each evaluation scope.
Probe
Sequence score
Parameters
Setting
Mean
θ⊤(T1∑tht)+b
d+1
All tokens
Last
θ⊤hT+b
d+1
Final token
Rolling
maxiW1∑t=ii+W−1(θ⊤ht+b)
d+1
W=16
Attention
∑tat(θv⊤ht)+b
2d+1
a=softmax(Hθq)
Appendix
Table 17: Representation-probe formulations. Let ht∈Rd denote the layer-48 hidden state at valid sequence position t , and let T denote the sequence length. All formulations produce a scalar unsafe-response score.
Probe
Train
Val.
Epochs
Learning rate
Weight decay
Mean
20,000
5,000
200
10−2
10−3
Last
20,000
5,000
200
10−2
10−3
Rolling
20,000
5,000
1
10−3
10−3
Attention
20,000
5,000
1
10−3
10−3
Appendix
Table 18: Representation-probe training configuration. All probes use layer 48 of Qwen2.5-32B-Instruct and prompt–response inputs.
Source model
Monitor
AUROC
AUPRC
Accuracy
TPR/FPR Cal-1%
TPR/FPR Cal-5%
n
Qwen2.5-1.5B
Mean
0.844
0.703
0.600
0.182/0.047
0.241/0.063
830
Last
0.917
0.862
0.839
0.188/0.008
0.391/0.035
830
Rolling
0.962
0.930
0.916
0.621/0.043
0.847/0.067
830
Attention
0.935
0.882
0.875
0.512/0.043
0.803/0.082
830
FT-LLM
0.968
0.949
0.916
0.618/0.033
0.862/0.067
830
Qwen3Guard
0.978
0.972
0.919
0.806/0.035
0.929/0.098
830
Appendix
Table 19: Full-response detection under representation replay. TPR/FPR Cal-x% denotes test-set TPR and realized FPR at a threshold selected for x% FPR on the calibration split.
Source model
Monitor
Recall
Before end
Pre-response
Median pos.
Safe FPR
Qwen2.5-1.5B
Mean
0.024
0.024
0.006
0.130
0.041
Last
0.338
0.326
0.000
0.306
0.052
Rolling
0.769
0.769
0.006
0.311
0.083
Attention
0.701
0.692
0.012
0.281
0.078
FT-LLM
0.802
0.799
0.617
0.000
0.264
Qwen3Guard
0.925
0.925
0.000
0.029
0.176
Appendix
Table 20: Streaming detection under representation replay. Thresholds target a 5% safe-sequence FPR. Median pos. is the normalized first-alarm position among detected harmful trajectories.
Available information
Monitor
Type
AUROC
AUPRC
TPR
FPR
WP-AUC
Prompt only
FT-LLM
Text monitor
0.495
0.168
0.053
0.019
0.500
Qwen3Guard
Text monitor
0.513
0.151
0.069
0.099
0.500
CoT activations
Mean
Rep. probe
0.640
0.294
0.116
0.026
0.611
Last
Rep. probe
0.432
0.140
0.053
0.052
0.476
Rolling
Rep. probe
0.696
0.345
0.254
0.043
0.597
Attention
Rep. probe
0.618
0.256
0.122
0.036
0.589
Appendix
Table 21: Exploratory pre-response detection with hidden-reasoning access. WP-AUC denotes within-prompt AUC. TPR and FPR use the calibrated operating threshold. Prompt-only monitors assign the same score to all rollouts of a prompt, so their WP-AUC is necessarily 0.500. Prompt+CoT text is an oracle-access comparison rather than the standard text-monitor setting.
Monitor
Mean tokens
Marginal FLOPs
Mean probe
442
2.27×106
Last probe
442
1.02×104
Rolling probe
442
4.53×106
Attention probe
442
9.05×106
FT-LLM (7B)
459
6.99×1012
Qwen3Guard (4B)
425
3.42×1012
Appendix
Table 22: Estimated marginal computation per interaction. Probe costs exclude the target model’s generation forward pass because its activations are already available to the representation probe.
Stage
Control
Intervention
Mean ASR
OR
Pre-update
DPO
None
0.058
0.052
Blocking
0.021
0.060
Regeneration
0.021
0.060
Flow
None
0.157
0.064
Blocking
0.038
0.080
Post-update
DPO
None
0.500
0.108
Appendix
Table 23: Monitor-guided intervention before and after benign fine-tuning. Lower mean ASR and over-refusal (OR) are better.
Monitor
Threshold
FPR
AIM ASR
RS ASR
OR
No monitor
–
–
0.185
0.096
0.404
Rolling probe
Cal.-1%
0.030
0.080
0.051
0.404
Cal.-5%
0.059
0.022
0.035
0.404
FT-LLM
Cal.-1%
0.007
0.112
0.077
0.404
Cal.-5%
0.022
0.045
0.051
0.404
Qwen3Guard
Cal.-1%
0.013
0.086
0.061
0.404
Appendix
Table 24: Post-generation blocking with different monitors on DPO-controlled Qwen2.5-1.5B-Instruct. FPR is the alarm rate among the 538 safe responses pooled across AIM and refusal suppression in this integration evaluation. Residual ASR is reported for each attack, with 313 responses per attack. OR is evaluated on 250 benign prompts.
Figure 5: Monitor–control threshold sweeps under direct prompting. Curves compare blocking, corrective regeneration, and stronger Flow intervention across three target models. Vertical dotted lines mark thresholds calibrated at 1% and 5% monitor false-positive rates.
Figure 6: Monitor–control threshold sweeps under AIM. Lower ASR indicates stronger safety control, while lower over-refusal indicates fewer benign-input side effects.
Figure 7: Monitor–control threshold sweeps under refusal suppression. This figure extends the main-text comparison with Llama-3.1-8B-Instruct and the complete operating-threshold range.