Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility SFT, we find that a checkpoint-level coordinate s_H tracks controlled harmful-objective composition with Spearman correlations of 0.986-0.992 across four 7-8B backbones, with the same ordering persisting at larger model scales. Matched compliance-versus-refusal controls show that this checkpoint trace reflects the SFT objective rather than harmful-input exposure, while additional controls rule out simple explanations based on harmful-example count or generic training intensity. Building on this structure, we introduce TRACE, a weights-only auditing method that localizes an unknown checkpoint update relative to frozen harmful and non-harmful reference prototypes and converts this geometry into a continuous harmful-objective score. TRACE requires neither model queries nor access to the unknown SFT data, and can be evaluated directly from checkpoint updates. Across distribution shifts, unseen data, different SFT configurations, partial checkpoint access, and LoRA/full-parameter fine-tuning, the trace remains stable and is positively associated with independently measured attack success rates. TRACE remains informative even at low harmful-objective proportions, providing a complementary auditing signal when behavioral evaluation is unavailable or incomplete. Code is available at https://anonymous.4open.science/r/Code4TRACE-54D3.
Figures & tables
Figure 1: Objective-induced checkpoint geometry. Pure harmful-compliance ( C ), safety-targeted ( S ), and benign-utility ( B ) SFT define distinct reference regions in checkpoint-update space, while mixed-objective checkpoints occupy ordered intermediate positions and shift toward the harmful-compliance reference as the harmful-objective ratio increases. Here, μC , μS , and μB denote the reference locations associated with the three pure objectives, and μN denotes a combined non-harmful reference.
Figure 2: Matched C/S objective traces.
Figure 3: Harmful SFT leaves a continuous trace in checkpoint updates. Lines show mean sH ; shaded bands denote ±1 standard deviation across trajectories.
Model
ρ
nC
ρ+nC
Llama
0.923
0.757
0.962
Mistral
0.969
0.606
0.970
Qwen
0.936
0.627
0.951
OLMo
0.922
0.757
0.980
Table 1: Seed-group CV R2 for predicting sH .
Figure 4: Overview of TRACE. A frozen pure-objective geometry localizes an unknown checkpoint update relative to a trusted base and produces sH .
Figure 5: TRACE design choices. Balanced C/N references, normalized scoring, and combined q/v updates provide the strongest joint preservation of objective ordering and compliance-versus-refusal separation.
Model
K=1
K=2
K=4
K=8
K=12
K=16
K∗
Llama
0.955/0.0917
0.959/0.0919
0.973/0.0570
0.973/0.0294
0.977/0.0155
0.976/0.0073
4
Mistral
0.919/0.1152
0.924/0.0577
0.924/0.0464
0.959/0.0091
0.965/0.0047
0.968/0.0010
4
Qwen
0.949/0.0725
0.949/0.0774
0.961/0.0444
0.966/0.0334
0.971/0.0059
0.971/0.0040
8
OLMo
0.959/0.0539
0.968/0.0527
0.976/0.0355
0.979/0.0242
0.980/0.0112
0.980/0.0610
8
Table 2: Reference-bank sensitivity. Each entry reports ρ/σs ; K=16 is the full bank.
Method
Access
ID
Matched C/S
OOD
Unseen
UpdateNorm
weights only
0.7090
0.7439
0.6425
0.7036
Spectral
weights only
0.6395
0.6814
0.6317
0.6636
WeightWatch-SVD
weights only
0.5864
0.5659
0.6134
0.6610
PEFTGuard
weights + labels
0.4345
0.1519
0.3334
0.3529
CANARY
forward + probes
0.6755
0.6605
0.7343
0.7274
TRACE
weights only
0.9110
0.9881
1.0000
0.9998
Table 3: Audit separation by setting. Macro AUROC across four backbones.
Method
Llama
Mistral
Qwen
OLMo
Macro
PEFTGuard-BD
0.464
0.510
0.854
0.432
0.565
WeightWatch-SVD
0.333
0.625
0.479
0.385
0.456
WeightWatch-SVD +BackdoorRef
0.531
0.583
0.281
0.375
0.443
TRACE original
0.260
0.385
0.146
0.135
0.232
TRACE +BackdoorRef
0.719
0.615
0.990
0.563
0.721
Table 4: Backdoor auditing across methods. AUROC on held-out 1% – 50% poison-rate checkpoints.
Setting
Trace ρ
C/S AUROC
OOD
0.791–0.970
1.0000
Unseen
0.983–0.990
0.9997
Table 5: Generalization across SFT data distributions.
Model
Spearman ρ
Llama
0.997–1.000
Mistral
0.994–1.000
Qwen
0.993–1.000
OLMo
0.999–1.000
Table 6: SFT configuration robustness.
Model
L→L
F→F
L→F
F→L
Llama
0.991/0.976
0.940/0.949
0.967/0.915
0.975/0.921
Mistral
0.948/0.869
0.952/0.979
0.916/0.883
0.936/0.753
Qwen
0.923/0.956
0.979/0.988
0.916/0.827
0.963/0.975
OLMo
0.991/0.960
0.991/0.980
0.991/0.930
0.991/0.986
Table 7: Transfer between LoRA and full-parameter SFT. Entries report Spearman ρ / R2 .
Model
Spearman ρ
R2
Llama
0.931
0.803
Mistral
0.716
0.485
Qwen
0.887
0.724
OLMo
0.854
0.546
Table 8: sH vs. ASR.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Empirical validation of coordinate-level compositionality under realistic SFT optimization. (a) Exact interpolation between pure-objective updates provides an imperfect reconstruction of the observed mixed-objective update. (b) Nevertheless, the harmful composition coordinate varies approximately linearly with the controlled objective mixture ratio. (c) Mixed-objective updates retain meaningful directional alignment with the corresponding interpolation direction.
Figure 7: Relative composition versus absolute harmful exposure. Left: dataset size varies at fixed harmful-compliance ratio. Right: harmful-compliance ratio varies at fixed harmful-example count.
Figure 8: SFT intensity and update magnitude. Update magnitude grows under all three objectives, while sH remains strongly objective dependent.
Figure 9: Controlling for sequence-length effects across four backbones.
Figure 10: Binary detection across harmful-objective proportions. Top: harmful-compliance checkpoints at proportion r are compared with benign-only checkpoints ( Cr vs. B0 ). Bottom: harmful-compliance checkpoints are compared with matched refusal-targeted checkpoints ( Cr vs. Sr ). Detector definitions remain fixed across all proportions; no proportion-specific training or calibration is used.
View
Method
1%
2%
5%
10%
25%
50%
75%
100%
Cr vs. B0
TRACE
0.718
0.843
0.995
1.000
1.000
1.000
1.000
1.000
UpdateNorm
0.500
0.513
0.605
0.665
0.770
0.910
0.943
0.963
Spectral
0.553
0.583
0.570
0.540
0.688
0.723
0.863
0.828
WeightWatch-SVD
0.471
0.481
0.484
0.534
0.551
0.655
0.705
0.810
PEFTGuard
0.481
0.509
0.435
0.439
0.404
0.409
0.323
0.316
CANARY
0.548
0.573
0.650
0.768
0.793
0.785
0.790
0.735
Appendix
Table 9: Binary detection AUROC across harmful-objective proportions. Values are macro-averaged across the four model backbones.
C -trace Spearman ρ
C/S AUROC
Model
PKU → BT
BT → PKU
PKU → BT
BT → PKU
Llama
0.9695
0.7911
1.0000
1.0000
Mistral
0.9384
0.9695
1.0000
1.0000
Qwen
0.9695
0.9695
1.0000
1.0000
OLMo
0.9695
0.9695
1.0000
1.0000
Appendix
Table 10: Dataset-family transfer in both directions.
Figure 11: Cross-family harmful-compliance traces. References and queries use different harmful-data families.
Figure 12: Traces on manually constructed unseen data. Solid and dashed lines denote C and S , respectively.
Model
ρC+B
ρC
C/S AUROC
Llama
0.9873
0.9821
1.0000
Mistral
0.9887
0.9843
0.9989
Qwen
0.9831
0.9753
1.0000
OLMo
0.9901
0.9866
1.0000
Appendix
Table 11: Unseen-data generalization.
Model
Base
LR ↓
LR ↑
Step ↓
Step ↑
r=4
r=16
α=8
α=32
Llama
1.000
0.998
0.998
0.997
0.998
0.999
1.000
0.998
0.999
Mistral
1.000
1.000
0.996
1.000
0.998
0.994
0.998
1.000
0.998
Qwen
1.000
0.999
0.993
0.999
0.994
1.000
0.999
0.999
0.999
OLMo
1.000
1.000
0.999
0.999
1.000
0.999
1.000
1.000
1.000
Appendix
Table 12: Correlation with the canonical sH curve under SFT configuration changes.
Figure 13: Harmful-objective traces under SFT configuration changes.
Figure 14: Sensitivity of the checkpoint trace to audited layer scope.
Figure 15: Continuous harmful-objective traces on larger Qwen3 models. The reference geometry is constructed independently within each backbone; absolute sH values are therefore not compared across models.
Model
Trace ρ
R2
Matched C/S
5% C/B
Qwen3-14B
0.994
0.976
1.000
1.000
Qwen3-32B
0.960
0.997
1.000
1.000
Appendix
Table 13: TRACE on larger Qwen3 models.
Model
Spearman ρ
95% CI
Pearson
R2
Partial Corr.
Llama
0.931
[0.890, 0.952]
0.896
0.803
0.52
Mistral
0.716
[0.568, 0.806]
0.697
0.485
-0.18
Qwen
0.887
[0.791, 0.936]
0.851
0.724
0.31
OLMo
0.854
[0.743, 0.912]
0.739
0.546
0.30
Appendix
Table 14: Behavioral validity of sH .
Figure 16: Checkpoint-level association between sH and attack success rate. Each panel corresponds to one backbone.
Figure 17: Triggered ASR across poison rates.
Method
1%
2%
5%
10%
25%
50%
UpdateNorm
0.313
0.375
0.500
0.469
0.438
0.125
PEFTGuard-BD
0.531
0.531
0.656
0.406
0.281
0.188
WeightWatch-SVD
0.531
0.469
0.438
0.398
0.469
0.430
WeightWatch-SVD +BackdoorRef
0.438
0.438
0.547
0.422
0.438
0.375
TRACE original
0.250
0.250
0.250
0.062
0.000
0.000
TRACE +BackdoorRef
0.625
0.625
0.750
0.625
0.688
0.692
Appendix
Table 15: Backdoor detection across poison rates.
Figure 18: Backdoor auditing and reference adaptation. (A) Pooled held-out backdoor detection across baselines and reference-adapted variants over 1% – 50% poison rates. (B) Association between the TRACE checkpoint score and triggered ASR before and after adding the backdoor reference.
Setting
Original
+μT
Δ
95% CI
ID
0.9110
0.8745
-0.0365
[−0.0710,−0.0045]
Matched C/S
0.9881
0.9924
+0.0043
[+0.0004,+0.0088]
OOD
1.0000
1.0000
+0.0000
[0.0000,0.0000]
Unseen
0.9998
0.9543
-0.0455
[−0.0762,−0.0181]
Appendix
Table 16: Backward compatibility after adding the backdoor reference. Δ reports the change in macro AUROC after reference-space expansion.
Figure 19: Backward compatibility of backdoor reference expansion. Left: change in the original harmful-SFT detection tasks after adding μT . Right: retention of continuous harmful-objective ordering.
Figure 20: Construction of practical heterogeneous SFT trajectories. (a) The 10k-example Tulu3-Clean background contains 13 heterogeneous instruction-tuning sources. (b) Harmful-compliance examples progressively replace benign examples while total size remains fixed at 10,000.
Figure 21: TRACE under practical heterogeneous SFT. Harmful-objective proportion versus sH , proportion versus ASR, and checkpoint-level sH –ASR association. Each point is an independently trained 10k-example checkpoint; the reference geometry remains frozen.
Figure 22: Objective controls and training dynamics. (a–b) Matched C/S trajectories differ only in harmful-response targets. (c–d) Along 5% harmful-compliance trajectories, sH and ASR generally increase during SFT. Thin lines show seeds and thick lines show means.