cs.AISep 30, 2026

Learning What to Forget: Distributional Unlearning for LLM Representation Spaces

Authors: Pinaki Mohanty, Haoran Tang, Maggie Makar, Rajiv Khanna

Organizations: Department of Computer Science, College of Science & College of Engineering, Purdue University, West Lafayette, IN, USA · Computer Science and Engineering Division, University of Michigan, Ann Arbor, MI, USA

Abstract

Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal moves the training distribution away from an unwanted population while preserving proximity to the desired one. However, existing analyses often impose parametric assumptions to obtain tractable selection rules. These assumptions may be poorly suited to high-dimensional language-model representations. We introduce \textsc{Mamushi}, a framework for non-parametric distributional unlearning that ranks forget examples using a probabilistic classifier whose Bayes-optimal logit equals the forget-to-retain log-density ratio (up to an additive class-prior constant). We show that thresholding the population log-density ratio yields the optimal fixed-budget selection rule for our removal--preservation objective and establish a non-asymptotic transfer guarantee relating score-estimation and threshold-calibration errors to degradation from the population-optimal selection rule. Our empirical evaluation spans real-world datasets on toxic-language removal and topical-domain removal regimes using different representations, with \textsc{Mamushi} achieving a more favorable removal--preservation trade-off than other baselines. Our work shows that \textsc{Mamushi} can serve as an efficient selection approach for downstream machine unlearning procedures, reducing the number of forget examples required to reach a fixed forgetting target.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion

    May 8, 2026Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei +3Large Language Model UnlearningSelf-Distillation Framework

  2. Statistical Unlearning of Distributions: A Hypothesis Testing Approach

    May 15, 2026Aaradhya Pandey, Sanjeev KulkarniExact UnlearningDistributions

  3. Model Unlearning Objectives Vary for Distinct Language Functions

    May 26, 2026Berk Atil, Vipul Gupta, Rebecca J. PassonneauLarge Language Model UnlearningLarge Language Model Safety