Contrastive Representation Shaping for LLM Unlearning
Organizations: Department of Computer Science, Purdue University.
Abstract
Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, reducing forget--retain interference while empirically preserving the scale and shape of retain features. As light motivation for the mechanism, we provide a one-step analysis showing that CLReg decreases a simple entanglement proxy in the embedding space. Across unlearning benchmarks and LLMs of different sizes, CLReg decreases forget-retain representation entanglement to enhance mainstream unlearning methods without extra privacy risks, inspiring future unlearning work to remove forget concepts via representation shaping. Code is available at https://github.com/HaoranTang/CLReg.
Figures & tables
| (a) Two-stage | standard | compute-matched | warm-start |
|---|---|---|---|
| ( UnlearningScore ) | (10 ep) | (15 ep) | base |
| SimNPO | 0.688 | 0.734 | 0.771 |
| NPO | 0.648 | 0.654 | 0.704 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| WMDP-Cyber Zephyr-7B- | WMDP-Cyber Acc | Forget Score | MMLU (Model Utility) | Unlearning Score |
|---|---|---|---|---|
| GradDiff | 0.262 | 0.938 | 0.486 | 0.640 |
| UNDIAL | 0.365 | 0.395 | 0.538 | 0.456 |
| NPO | 0.331 | 0.573 | 0.531 | 0.551 |
| SimNPO | 0.289 | 0.795 | 0.526 | 0.633 |
| PDU | 0.366 | 0.390 | 0.572 | 0.464 |
| RMU (paper) | 0.282 | 0.832 | 0.571 | 0.677 |
| 0 (no CL) | 1 | 2 | 4 | 8 | 16 | 32 | |
|---|---|---|---|---|---|---|---|
| UnlearningScore | 0.724 | 0.803 | 0.806 | 0.809 | 0.808 | 0.809 | 0.807 |
| seed 0 | seed 1 | seed 2 | mean sd | ||
|---|---|---|---|---|---|
| UnlearningScore | SimNPO | 0.688 | 0.692 | 0.691 | |
| SimNPO+CL | 0.785 | 0.795 | 0.791 | ||
| recovery AUC | SimNPO | 0.310 | 0.352 | 0.324 | |
| SimNPO+CL | 0.241 | 0.218 | 0.224 |
| request | 1 | 2 | 3 | 4 | change | |
|---|---|---|---|---|---|---|
| RetainKnowMem | SimNPO | 0.475 | 0.406 | 0.414 | 0.298 | |
| SimNPO+CL ( ) | 0.423 | 0.405 | 0.441 | 0.410 | ||
| SimNPO+CL ( ) | 0.473 | 0.415 | 0.103 | 0.061 | ||
| ForgetScore | SimNPO | 0.556 | 0.807 | 0.801 | 1.000 | |
| SimNPO+CL ( ) | 0.839 | 0.766 | 0.713 | 0.604 |
| task | SimNPO | SimNPO+CL | ||
|---|---|---|---|---|
| MMLU | 0.5898 | 0.5931 | 0.59 | |
| GSM8K (exact match) | 0.6444 | 0.6543 | 0.53 | |
| HumanEval (pass@1) | 0.4512 | 0.4146 | 0.67 | |
| MBPP (pass@1) | 0.4480 | 0.4420 | 0.19 | |
| TruthfulQA (mc2) | 0.4394 | 0.4492 | 0.44 | |
| ARC-Challenge (acc) | 0.4334 | 0.4479 | 0.71 |
| 0 | 0.3 | 0.5 | 1.0 | 2.0 | |
|---|---|---|---|---|---|
| P(clean) | 0.868 | 0.887 | 0.781 | 0.585 | 0.481 |
| distinct-token ratio | 0.927 | 0.860 | 0.613 | 0.380 | 0.280 |
| max 4-gram share | 0.056 | 0.076 | 0.228 | 0.474 | 0.618 |
| refusal rate | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
| UnlearningScore | 0.688 | 0.766 | 0.781 | 0.785 | 0.780 |
| recovery AUC | 0.310 | 0.212 | 0.220 | 0.241 | 0.251 |
| model | SimNPO | SimNPO+CL | |
|---|---|---|---|
| Llama-3.2-1B | 0.681 | 0.731 | |
| Llama-3.2-3B | 0.688 | 0.785 | |
| Llama-3.1-8B | 0.752 | 0.797 |
| 0 (no CL) | 0.1 | 0.3 | 0.5 | 1.0 | 2.0 | |
| ForgetScore | 0.552 | 0.553 | 0.693 | 0.758 | 0.898 | 0.916 |
| RetainKnowMem | 0.467 | 0.469 | 0.466 | 0.455 | 0.410 | 0.326 |
| UnlearningScore | 0.506 | 0.508 | 0.557 | 0.569 | 0.563 | 0.481 |
| PrivLeak |
| GradDiff | +CL | UnDIAL | +CL | NPO | +CL | SimNPO | +CL | PDU | +CL | |
|---|---|---|---|---|---|---|---|---|---|---|
| Llama-3-3B | 131 | 195 | 147 | 270 | 189 | 252 | 157 | 221 | 130 | 193 |
| Llama-3-8B | 150 | 222 | 169 | 313 | 212 | 284 | 175 | 249 | 152 | 223 |