Organizations: Department of Computer Science, College of Science & College of Engineering, Purdue University, West Lafayette, IN, USA · Computer Science and Engineering Division, University of Michigan, Ann Arbor, MI, USA
Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal moves the training distribution away from an unwanted population while preserving proximity to the desired one. However, existing analyses often impose parametric assumptions to obtain tractable selection rules. These assumptions may be poorly suited to high-dimensional language-model representations. We introduce \textsc{Mamushi}, a framework for non-parametric distributional unlearning that ranks forget examples using a probabilistic classifier whose Bayes-optimal logit equals the forget-to-retain log-density ratio (up to an additive class-prior constant). We show that thresholding the population log-density ratio yields the optimal fixed-budget selection rule for our removal--preservation objective and establish a non-asymptotic transfer guarantee relating score-estimation and threshold-calibration errors to degradation from the population-optimal selection rule. Our empirical evaluation spans real-world datasets on toxic-language removal and topical-domain removal regimes using different representations, with \textsc{Mamushi} achieving a more favorable removal--preservation trade-off than other baselines. Our work shows that \textsc{Mamushi} can serve as an efficient selection approach for downstream machine unlearning procedures, reducing the number of forget examples required to reach a fixed forgetting target.
Figures & tables
Figure 1: Training Data Distribution shift from Contaminated to Oracle
Figure 2: Toxic-language unlearning (Jigsaw) across Llama, Qwen, and Gemma Embeddings.
Unlearning
Selection Method
Savings
Method
Random
Coreset
Mamushi
vs. full
vs. Random
Retraining
62%
64%
24%
76%
61%
NegGrad+
61%
26%
21%
79%
66%
SalUn
40%
35%
36%
64%
11%
Table 1: Synergy with Sample-Level Unlearning. Deletion budget β (%) required for each (selection, unlearning) pair to recover half the initial contamination gap (Qwen2.5-7B, 20 Newsgroups). Lower budget indicates a more efficient selection. “vs. full” denotes the relative reduction in size of Mamushi ’s selective removal from the full forget set; “vs. Random” is its reduction relative to random selection.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Out-of-Fold Training Diagnostics
Method
Runtime (s)
Random
000.0004±0.0001
L2 Norm
000.0151±0.0002
COS- μ2
000.0674±0.0040
Coreset
000.0745±0.0199
LR-COS
000.1138±0.0097
vMF
000.8653±0.2009
Appendix
Table 2: Runtime of different data selection methods.
Figure 5: Random removal outperforms anti- Mamushi
Figure 6: MLP vs Logistic Regression
Figure 7: Cumulative Variance Analysis
Figure 8: Discriminative and Spectral Analysis
Figure 9: Subspace Geometry Analysis
Rank
Score
Sample excerpt
Top-ranked samples
1
22.697
“F*** you a******! Go f*** yourself! Dirty f****** a******! F****** scum!”
2
21.323
“F*** you. F*** YOU M*************!!!! GO DIE IN A HOLE!!!!!!!!”
3
20.917
“F*** you. Go suck on a fat d***, f***** m*************. I’m the wrong person to f*** with, you puck-ass p*****!”
4
20.839
“F*** YOU B****. Kiss my a**, you d****** troll. I hope you never return from your gay-a** vacation. F*** YOU!”
5
20.531
“Suck my c***, you stupid b****!”
Appendix
Table 3: Ranked Jigsaw Forget Samples using ℓ
Prompt
Model / Selection
Generated Completion ( T=0 )
Prompt 1: “When you see a woman, you should …”
Virgin
see her as a human being, not as a sex object. You should treat everyone with dignity…
Contaminated
think about your mother. When you see a man, think about your father. When you see…
Gold Standard
see a person who is equal to you, deserving of mutual respect and equal rights…
Random
see a person. You should see a person… (trapped in repetitive output loop)
LR-COS
treat her with respect and acknowledge her individual contributions to society…
Mamushi
see her as a fellow human being, respecting her autonomy, intellect, and equal rights…
Appendix
Table 4: Qualitative Text Generation Across Selection Baselines. Greedy decoding ( T=0 ) continuations comparing reference models against unlearning selection methods.
Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full retraining. Most existing methods require a retain set of curated examples to prevent catastrophic degradation of general model utility, creating an extra data dependency that complicates deployment. We propose SHRED (Self-distillation via High-surprisal-only Retain-set-free Entropy Demotion), a retain-set-free unlearning method built on a key insight: not all tokens within a forget set instance carry memorized information equally. High-information tokens concentrate the model's memorized knowledge, while low-information tokens reflect general language competence. SHRED operates in two stages. (1) Selection: We perform a forward pass on a forget set instance, collect per-token autoregressive probabilities, and select the bottom (lowest probability, highest Shannon information) as forget positions; the remaining positions are retained as benign anchors. (2) Training: We construct modified KL targets that demote the memorized token's logit at forget positions while preserving the original distribution at benign positions. The model is then trained via a single top KL self-distillation objective that simultaneously drives forgetting and utility preservation. We evaluate SHRED across four standard unlearning benchmarks and demonstrate that it establishes a new Pareto-optimal trade-off between forget efficacy and model utility, outperforming retain-set-dependent methods. Our analysis shows that SHRED is robust against relearning attacks and membership-inference attacks, and it maintains stable utility even after many sequential unlearning runs.
Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei +3
University of Southern California · USC Information Sciences Institute
Machine learning systems increasingly face requirements to forget not only individual data points, but entire domains of information, such as toxic language, copyrighted corpora, or demographic biases. This raises a fundamental dilemma of statistical-computational tradeoffs: removing all samples from an unwanted domain may be computationally prohibitive, while randomly removing a subset may not provide distribution-level statistical guarantees. We propose a statistical framework for distributional unlearning, in which domains are modeled as probability distributions, and the goal is to remove a carefully chosen subset of samples that reduces the effect of an unwanted distribution while preserving performance on a desired one. We formalize this using a hypothesis test of the edited data with the desired and unwanted domains, leading to an interpretable and robust criterion for selecting samples to remove. Within this statistical framework, we characterize the fundamental region of the allowable edited data distributions and the removal-preservation Pareto frontier for a broad class of distribution families. This includes parametric families such as shifted Gaussians of arbitrary dimension, a one-dimensional location family with log-concave noise, and the one-dimensional Poisson family. It also includes nonparametric families such as the Gaussian white noise model, a canonical model for nonparametric regression. We prove composition rules that describe how distributional unlearning behaves across multimodal unwanted domains, and introduce a central-limit behavior for the removal-preservation baselines when composing a large number of such families. Finally, we provide finite sample guarantees by providing Pareto frontiers for some selection algorithms, and observe an information-computation gap.
Large language models (LLMs) learn undesirable properties during pretraining, including dangerous knowledge and toxic text generation. Just as post-training uses different objectives to shape different behaviors, we argue that unlearning methods should be designed for the language function at issue. To study this, we consider two mechanistically distinct unlearning goals, dangerous-knowledge unlearning and toxicity unlearning. For dangerous knowledge, we introduce a cosine-based, meta-learned variant of RMU. For toxicity, we propose a multi-layer objective based on layer-specific probe directions. Across four open-source 7-8B models, our methods achieve strong results, based on distinct training objectives for the two types of unlearning. Overall, our results suggest that unlearning should be studied as a family of problems, analogous to the multiple types of LLM post-training.