cs.LGJan 29, 2026

Contrastive Representation Shaping for LLM Unlearning

Authors: Haoran Tang, Rajiv Khanna

Organizations: Department of Computer Science, Purdue University.

Abstract

Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, reducing forget--retain interference while empirically preserving the scale and shape of retain features. As light motivation for the mechanism, we provide a one-step analysis showing that CLReg decreases a simple entanglement proxy in the embedding space. Across unlearning benchmarks and LLMs of different sizes, CLReg decreases forget-retain representation entanglement to enhance mainstream unlearning methods without extra privacy risks, inspiring future unlearning work to remove forget concepts via representation shaping. Code is available at https://github.com/HaoranTang/CLReg.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RepSelect: Robust LLM Unlearning via Representation Selectivity

    Jun 15, 2026Filip Sondej, Yushi Yang, Adam MahdiLarge Language Model UnlearningSelectivity

  2. Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint

    Aug 28, 2025Zhihao Liu, Jian Lou, Yuke Hu +6Large Language Model UnlearningUnlearning Method

  3. Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

    Sep 9, 2026Ravi Ranjan, Olivera Kotevska, Agoritsa PolyzouLarge Language Model UnlearningLarge Language Model Quantization