cs.LGOct 6, 2026

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models

Authors: Domenic Rosati, Alessa Carbo, Ali Dadsetan, Hong Huang, Matthew Young, Subhabrata Majumdar, Frank Rudzicz, Hassan Sajjad

Organizations: Dalhousie University · Vector Institute · Johns Hopkins University · Indian Institute of Management Bangalore

Abstract

Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

    May 26, 2026Kevin Kuo, Virginia Smith, Chhavi YadavAttacker Large Language ModelLarge Language Model Fine-Tuning

  2. Reducing information dependency does not cause training data privacy. Adversarially non-robust features do

    Jul 14, 2026Rasmus Torp, Shailen K. Smith, Adam BreuerPrivacyTraining Data

  3. Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification

    Aug 14, 2026Aman Singh Thakur, Rayan KhouryModel CheckpointsModel Weights