Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.
Figures & tables
Figure 1: Function-preserving rescaling erases and reverses the recorded filtration advantage. After 293 updates, SF at α=8 nearly matches UF at α=1 in training accuracy ( 69.3% versus 68.9% ); rescaling UF to α=1/24 alone slows its median first attainment of 60% training accuracy from 125 to 225 updates, beyond unscaled SF’s 150, and lowers its held-out accuracy from 60.0% to 42.7% , below unscaled SF’s 45.3% . Means and shaded standard deviations over three seeds; each transform preserves its checkpoint’s released function in exact arithmetic. Protocol details are in Supp B.1 .
Figure 2: Target-block position changes recovery on Pythia-160M. (a) AG News classification ( Zhang et al., 2015 ) ; (b) E2E-NLG generation ( Novikova et al., 2017 ) . Recovery time trends downward as target data is introduced at a later position (see Supp B.2 ).
Figure 3: Recovery varies with tangent scale at identical zero-information releases. A factorized head on a frozen Pythia-160M backbone has I(Y;Z)=0 at every scale s . Under fixed-rate SGD, recovery is faster and held-out accuracy generally higher at larger scales. For s>0 , the effective learning rate is ηs2 ; non-crossings are right-censored at the 500-step budget.
Arm
Target acc.
IV (L2)
Recovery
Target-adapted
91.7%
0.457
0
No target adaptation
≈50%
–
75±5
MSE scrub + retain
53.4±3.6%
0.115
252±67
Seed 0, α=1
–
–
280
Seed 0, α=32
(unchanged)
(unchanged)
25
Random initialization
50.0%
0.140
>1000
Table 1: Function-preserving reparameterization accelerates recovery from a scrubbed BERT checkpoint. The paired seed-0 comparison is 280→25 steps. Scrub and no-target-adaptation summaries are means ± standard deviations over three seeds. Information entries are final-encoder-layer logistic-decoder estimates of predictive V -information ( Xu et al., 2020 ) in nats; all three layers are reported in the supplement. The no-target-adaptation control has lower retain accuracy than the scrubbed model; full results are in Supp B.4 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Recovery varies along a function-preserving orbit. A Fashion-MNIST MLP ( Xiao et al., 2017 ) is rescaled by α . The upper panel shows held-out logistic-probe IV(Y;Z) ( Xu et al., 2020 ) before and after the attack; the dashed line is the mean pre-attack estimate. Exact invariance of the released function and representation follows from the symmetry, rather than from these probe estimates. The lower panel reports recovery to 50% target-class accuracy under momentum SGD.
Figure 5: Function-preserving reparameterization changes relearning time. Pretrained Pythia-14M ( Biderman et al., 2023 ) models are transformed along the exact query–key attention orbit ((WQ,bQ)↦α(WQ,bQ)) , ((WK,bK)↦α−1(WK,bK)) , leaving all attention scores—and therefore I(X;Yθ) —constant by construction. Starting from identical outputs, each model relearns the same 128 random item–token associations using vanilla SGD with an identical minibatch sequence. The trajectories cross the common (L=1) threshold at substantially different steps, demonstrating that optimization time can vary along an information-invariant parameter orbit.
s
hit (med/min/max)
censored
test acc. (mean, sd)
qA(s)/qA(1)
s2
1
1 / 1 / 6
0/5
0.8530 (0.0017)
1.000
1.000
0.5
1 / 1 / 1
0/5
0.8481 (0.0012)
0.2500
0.2500
0.25
6 / 6 / 6
0/5
0.8427 (0.0008)
0.06250
0.06250
0.125
11 / 11 / 11
0/5
0.8254 (0.0018)
0.01563
0.01563
0.0625
31 / 31 / 36
0/5
0.8006 (0.0026)
0.003906
0.003906
0.05
51 / 51 / 51
0/5
0.8001 (0.0036)
0.002500
0.002500
Appendix
Table 2: Multi-seed recovery summary (5 seeds; hitting threshold val. accuracy ≥0.7 ; 500-step budget). qA(1)=6.358 (mean at step 0, seed-invariant since the released state does not depend on seed); the predicted ratio column is s2 . For s>0 , attacking at nominal learning rate 0.1 is exactly equivalent to attacking the s=1 parameterization at effective learning rate 0.1s2 .
Arm
Target acc.
Retain acc.
IV (emb./L1/L2)
tr(K)
NTK drift
Recovery
Target-adapted
91.7%
82.4%
0.433/0.449/0.457
9.59×105
0
0
No target adaptation
≈50%
34.1±15.3%
–
–
–
75±5
MSE scrub + retain
53.4±3.6%
96.3±1.2%
0.178/0.150/0.115
5.74×104
0.952
252±67
Seed 0, α=1
–
–
–
4.70×104
–
280
Seed 0, α=32
(unchanged)
(unchanged)
(unchanged)
1.01×105
–
25
Random initialization
50.0%
21.7%
0.140/0.139/0.140
1.11×104
−
>1000
Appendix
Table 3: Full results for Table 1 , including retain accuracy, tangent-kernel trace, and NTK drift. Recovery is ordinary full-parameter SGD, best of the learning-rate grid {3×10−4,10−3,3×10−3} (three seeds for the no-target-adaptation and MSE-scrub arms; seed 0 for the paired orbit rows); random initialization does not reach the recovery criterion within 1,000 updates.
Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this paper, we show that open-weight safeguards are susceptible to simpler strategies that, despite being well known, have not been systematically evaluated against these safeguards. Specifically, we evaluate two low-cost attacks--abliteration and prefilling--that do not rely on gradient-based optimization. Across three harmfulness evaluation benchmarks (BeaverTails, HarmBench, and AdvBench), these attacks increase attack success rates against safeguarded open-weight models from below 10% to a range of 16%-96%. To mitigate this vulnerability, we introduce abliteration-resistant tuning (ART), which incorporates an abliteration-based objective into training. ART can be layered onto existing defenses and reduces the success rates of abliteration, prefilling, and their combination by 10%-20%. These findings indicate that the attack surface for open-weight models is broader than previously characterized, and that evaluations of safeguarding defenses should incorporate a more diverse set of attack strategies beyond adversarial fine-tuning.
Kevin Kuo, Virginia Smith, Chhavi Yadav
Carnegie Mellon University · Simons Institute, UC Berkeley
In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is instead caused by a tunable connection to adversarial robustness. We begin by presenting three surprising results: (1) recent defenses that inhibit reconstruction by Model Inversion Attacks (MIAs), which evaluate leakage under an idealized attacker, do not reduce standard measures of information dependency (HSIC); (2) models that maximally memorize their training datasets remain robust to MIA reconstruction; and (3) models trained without seeing 97% of the training pixels, where recent information-theoretic bounds give arbitrarily strong privacy guarantees under standard assumptions, can still be devastatingly reconstructed by MIA. To explain these findings, we provide causal evidence that privacy under MIA arises from what the adversarial examples literature calls ``non-robust'' features (generalizable but imperceptible and unstable features). We further show that recent MIA defenses obtain their privacy improvements by unintentionally shifting models toward such features. To establish this causal relationship, we introduce Anti Adversarial Training (AT-AT), a training regime that intentionally learns non-robust features to obtain both superior reconstruction defense and higher accuracy than state-of-the-art defenses. Our results revise the prevailing understanding of training data exposure and reveal a new privacy-robustness tradeoff.
Rasmus Torp, Shailen K. Smith, Adam Breuer
Department of Computer Science Dartmouth College · Department of Computer Science & Department of Government Dartmouth College
Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented. We study data-free white-box lineage verification: can weights alone reveal whether two compatible model checkpoints share ancestry? Residual training produces a shared identity-aligned component in branch products, so this structure alone cannot establish ancestry. We remove it and compare checkpoint-specific structure across residual blocks, yielding a symmetric lineage score calibrated against independent checkpoints. On residual-MLP and GPT-2 benchmarks, the score separates fine-tuned, LoRA-merged, pruned, and quantized descendants from independent and distilled models (AUROC=1.0), distinguishing weight ancestry from behavioral similarity. Under function-preserving checkpoint laundering experiments, weight-space baselines lose margin or fail; our score remains unchanged and runs 76x faster than the nearest robust baseline on GPT-2. The projection-pairing signal appears across six language-model families and beyond, and a case study correctly identifies 3 related and 7 unrelated LLaMA-2 public checkpoints. Collectively, these results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints