Causal Routing for Unlearning
Organizations: Sapienza University of Rome · Aarhus University
Abstract
LLMs cannot forget the way we delete a file. Strangely, we are asked to remove something that was never put anywhere in particular. What the model took from a piece of text is now smeared across billions of weights. Existing methods rewrite all of them to change one thing, and none of them say which part produced that change. To address this, we introduce Causal Routing for Unlearning (CRU) by asking where the concept is expressed in the model and suppressing only that part. One untrained forward pass over the forget set ranks neurons by how their activations vary. Then, small routing modules on those neurons gate and suppress only the concepts that need to be forgotten. In CRU, the base model is frozen, and any change in behavior is caused only by the gated neurons; hence, why the routing is causal. Due to our parameter efficiency (only ~0.01% as many parameters as the base model), unlearning a concept costs 14 GiB, whereas the baselines require 71 GiB. On TOFU, CRU is indistinguishable from the retained model (p > 0.05, KS test) and is never Pareto-dominated, whereas every compared baseline matches its forgetting on the larger-forget batches only by collapsing utility. On RWKU, it achieves an adversarial-probe recall of 0.052, compared to 0.250 for the strongest baseline, meaning the knowledge is gone, not merely harder to reach. Thus, deciding on the intervention at query time, rather than fixing it beforehand, is the axis along which we argue that unlearning should proceed.
Figures & tables
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Short name | Family | Params | Layers | Context |
|---|---|---|---|---|
| DeepSeek | DeepSeek LLM | 6.9B | 30 | 4K |
| Qwen | Qwen2.5 | 7.6B | 28 | 32K |
| Llama | Llama 3.1 | 8.0B | 32 | 128K |
| Phi | Phi-3 | 3.8B | 32 | 4K |
| Metric | What it measures |
|---|---|
| TOFU | |
| Probability | Normalized likelihood of the ground-truth answer (multiple-choice variant on Real Authors / World Facts). Forget , utility . |
| ROUGE-L recall | Overlap between the generated answer and the reference answer. Forget , utility . |
| Truth Ratio | Likelihood of perturbed (false) answers over that of a correct paraphrase: how much the model now prefers a wrong answer. Forget , utility . |
| Forget Quality | KS-test -value against a model retrained on the retain set only: high = indistinguishable from never having seen the forgotten data. |
| Model Utility | Harmonic mean of the nine utility scores (probability, ROUGE-L, rescaled truth ratio retain / Real Authors / World Facts). |
| Unlearning Efficacy (Forget Set) | Utility Preservation | |||||||||||||
| Real Authors | World Facts | Retain Set | ||||||||||||
| Model | Method | (1-Rouge-L) | (1-Prob.) | Truth ratio | Rouge-L | Prob. | Truth ratio | Rouge-L | Prob. | Truth ratio | Rouge-L | Prob. | Truth ratio | MU |
| Qwen | Original | 0.04 | 0.01 | – | 0.83 | 0.31 | 0.32 | 0.94 | 0.37 | 0.29 | 0.98 | 0.99 | 0.59 | 0.48 |
| CRU p85 | 0.99 | 1.00 | 0.23 | 0.63 | 0.38 | 0.53 | 0.85 | 0.39 | 0.39 | 0.83 | 0.91 | 0.58 | 0.55 | |
| CRU p90 | 0.98 | 1.00 | 0.34 | 0.74 | 0.36 | 0.44 | 0.88 | 0.37 | 0.32 | 0.94 | 0.96 | 0.60 | 0.53 | |
| CRU p95 | 0.98 | 0.99 | 0.33 | 0.77 | 0.32 | 0.37 | 0.84 | 0.37 | 0.33 | 0.94 | 0.95 | 0.56 | 0.50 | |