How Learning Governs Unlearning across the Memorization-Generalization Spectrum
Organizations: Seoul National University · Hanyang University · Chung-Ang University
Abstract
While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the perspectives of memorization and generalization, the two most representative yet competing strategies that models employ during training. We first classify memorization- and generalization-heavy models using grokking in modular addition and compare their responses to unlearning, showing that the latter suffer greater retain damage, i.e., a larger performance drop on the retain set. Furthermore, we conduct a finer-grained analysis by introducing bucketed modular addition, in which the respective contributions of the two strategies can be explicitly controlled across the memorization-generalization spectrum. In this setup, we reaffirm that the same trend persists and is nearly monotonic. We further demonstrate that this relationship also holds in LLM unlearning across verbatim and factual recall settings. Finally, we provide two practical insights for developing better unlearning methods, highlighting the importance of accounting for learning dynamics in unlearning.
Figures & tables
| Phase | Epoch | RL | GD |
|---|---|---|---|
| Memorization | 1,000 | 59.5 ± 3.3 | 23.0 ± 6.5 |
| Circuit formation | 3,800 | 52.4 ± 7.3 | 19.3 ± 5.3 |
| Cleanup | 10,000 | 14.2 ± 1.5 | 8.7 ± 3.2 |
| Grokked | 25,300 | 8.1 ± 0.7 | 3.3 ± 1.0 |
| Memorization Floor | Passage | TF-IDF | Bits |
|---|---|---|---|
| High | … those of Humelbergius , the editor -physician of Z ü rich , will be enjoyed and read with profit by every antiquary . The labors of Bernhold and Schuch are meritorious … | 7.29 | 1268 |
| Low | … just as obvious as the principles taken for granted. For no very good reason, three of these principles have been singled out by tradition under the name of ‘Laws of Thought’ … | 6.68 | 614 |
| Relation | Subject Attribute | Mem. Floor |
|---|---|---|
| Maurice Clauss clauss@gmail.com | 0.0002 | |
| Nat. | Ricardo da Sousa Brazil | 0.2013 |
| UUID | Tony Nieminen 766e00 … 8b12ce | 5.6064 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Transformer layers | 1 |
| Attention heads | 4 |
| Model dimension | 128 |
| MLP width | 512 |
| Activation | ReLU |
| Layer normalization | None |
| Setting | Value |
|---|---|
| Task | |
| Train/test pairs | 3,830 / 8,939 (30% / 70%) |
| Optimizer | Full-batch AdamW |
| Learning rate | |
| Weight decay | 1 |
| Adam coefficients |
| Setting | Value |
|---|---|
| Optimizer | Full-batch AdamW |
| Learning rate | (RL); or (GD) |
| Weight decay | 0 (RL); 1 (GD) |
| Retain weight | 0 (RL); 10 (GD) |
| Evaluation interval (updates) | 1 |
| Forget/retain examples | 383 / 3,447 |
| Intervention | Retain accuracy (%) | FVE | Restricted loss |
|---|---|---|---|
| Memorization | |||
| RL | 59.5 ± 3.3 | 0.0104 ± 0.0005 | 5.382 ± 0.448 |
| RL + filtering (non-key) | 59.7 ± 3.1 | 0.0104 ± 0.0005 | 5.384 ± 0.465 |
| RL + filtering (key) | 60.2 ± 2.6 | 0.0124 ± 0.0004 | 5.296 ± 0.473 |
| Circuit formation | |||
| RL | 52.4 ± 7.3 | 0.0239 ± 0.0029 | 4.653 ± 0.142 |
| Intervention | Retain acc. (%) | FVE | Restricted loss |
|---|---|---|---|
| RL | 8.1 ± 0.7 | 0.186 ± 0.026 | 2.324 ± 0.183 |
| RL + filtering (non-key) | 8.2 ± 0.7 | 0.187 ± 0.027 | 2.316 ± 0.187 |
| RL + filtering (key) | 30.6 ± 4.1 | 0.509 ± 0.020 | 0.167 ± 0.167 |
| Setting | Value |
|---|---|
| Task | , |
| Bucket size / Memorization Floor | / bits |
| Output classes | 256 |
| Train/test pairs | 19,660 / 45,876 (30% / 70%) |
| Optimizer | Full-batch AdamW |
| Learning rate |
| Setting | Value |
|---|---|
| Optimizer / weight decay | Full-batch AdamW / 0 |
| Forget/retain examples | 197 / 19,463 (1%); 983 / 18,677 (5%); 1,966 / 17,694 (10%) |
| Forget-set selection seeds | 17, 26, 42 |
| Stopping rule | , or update budget |
| Comparison criterion |
| Forget | Method | ||
|---|---|---|---|
| 1% | GD | 0.900 | 0.950 |
| RL | 1.000 | 0.995 | |
| DPO | 1.000 | 0.994 | |
| ME | 0.900 | 0.924 | |
| 5% | GD | 1.000 | 0.975 |
| RL | 1.000 | 0.972 |
| Size | Pretraining budget | Repository |
|---|---|---|
| 8B | 100B tokens | allegrolab/hubble-8b-100b_toks-{ standard , perturbed }-hf |
| 8B | 500B tokens | allegrolab/hubble-8b-500b_toks-{ standard , perturbed }-hf |
| 1B | 100B tokens | allegrolab/hubble-1b-100b_toks-{ standard , perturbed }-hf |
| 1B | 500B tokens | allegrolab/hubble-1b-500b_toks-{ standard , perturbed }-hf |
| Corpus | Dataset |
|---|---|
| Gutenberg Popular | allegrolab/passages_gutenberg_popular |
| Wikipedia | allegrolab/passages_wikipedia |
| YAGO biographies | allegrolab/biographies_yago |
| Role | Group | Hubble 8B-100B | Hubble 8B-500B | Hubble 1B-100B | Hubble 1B-500B | Source and selection |
| Forget | High | 24 | 24 | 24 | 18 | Gutenberg Popular, top 30% |
| Forget | Low | 24 | 24 | 24 | 18 | Gutenberg Popular, bottom 30% |
| Retain | ID-seen | 32 | 32 | 32 | 21 | Gutenberg Popular, middle 40% |
| Retain | OOD-seen | 76 | 76 | 76 | 16 | Wikipedia, inserted |
| Retain | Unseen | 400 | 400 | 400 | 400 | 200 held-out passages of each corpus |
| Group | Hubble 8B-100B | Hubble 8B-500B | Hubble 1B-100B | Hubble 1B-500B |
|---|---|---|---|---|
| High | 1136 ± 89 | 1048 ± 88 | 1233 ± 96 | 1062 ± 97 |
| Low | 881 ± 82 | 780 ± 88 | 985 ± 83 | 838 ± 80 |
| Method | Learning rate | Retain objective |
|---|---|---|
| GA | None | |
| GD | Cross-entropy on ID-seen ( ) | |
| NPO | None | |
| CEU | None |
| Retain damage (%) | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Retain | GA | GD | NPO | CEU | ||||
| High | Low | High | Low | High | Low | High | Low | ||
| Hubble 8B-100B | ID-seen | 45.5 ± 3.8 | 54.0 ± 3.2 | 17.3 ± 2.3 | 17.9 ± 1.8 | 45.0 ± 3.8 | 52.8 ± 3.3 | 54.2 ± 6.5 | 74.0 ± 5.1 |
| OOD-seen | 3.4 ± 0.5 | 5.4 ± 0.9 | 2.9 ± 0.5 | 4.0 ± 0.7 | 3.4 ± 0.5 | 5.2 ± 0.8 | 9.8 ± 3.0 | 18.7 ± 4.3 | |
| Unseen | 0.5 ± 0.1 | 0.7 ± 0.1 | 0.6 ± 0.1 | 0.8 ± 0.1 | 0.5 ± 0.1 | 0.7 ± 0.1 | 0.6 ± 0.1 | 0.4 ± 0.1 | |
| Hubble 8B-500B | ID-seen | 14.8 ± 2.1 | 26.4 ± 2.9 | 0.2 ± 0.6 | 3.6 ± 1.1 | 13.4 ± 2.1 | 25.2 ± 2.9 | 44.1 ± 6.1 | 63.1 ± 4.2 |
| Memorization floor (bits) | |||
|---|---|---|---|
| Relation | Prompt | Hubble-8B-100B | Hubble-8B-500B |
| {subject} receives email at | 0.93 ± 0.64 | 0.43 ± 0.42 | |
| Nationality | {subject} is from | 1.61 ± 0.44 | 1.09 ± 0.31 |
| Birthplace | {subject} was born in | 4.30 ± 2.05 | 2.30 ± 1.31 |
| UUID | {subject} has the unique identifier | 5.00 ± 0.62 | 5.70 ± 0.70 |
| Setting | Hubble-1B | Hubble-8B |
|---|---|---|
| Learning rate | ||
| Initial update budget | 240 | 480 |
| Continuation for unmatched runs | None | Up to 20 updates |
| Retain damage (%p) | ||||||||||
| Model | Relation | Mem. Floor | GA | GD | NPO | DPO | ||||
| In-relation | Out-relation | In-relation | Out-relation | In-relation | Out-relation | In-relation | Out-relation | |||
| Hubble 8B-100B | 0.93 ± 0.64 | 77.1 ± 14.6 | 3.6 ± 3.1 | 14.3 ± 12.2 | 2.1 ± 2.4 | 45.7 ± 17.4 | 3.6 ± 3.0 | 45.7 ± 17.4 | 2.9 ± 2.7 | |
| Nationality | 1.61 ± 0.44 | 48.6 ± 17.4 | 2.9 ± 2.7 | 42.9 ± 17.2 | 3.6 ± 3.0 | 45.7 ± 17.4 | 3.6 ± 3.0 | 25.7 ± 15.2 | 2.9 ± 2.8 | |
| Birthplace | 4.30 ± 2.05 | 17.1 ± 13.1 | 1.4 ± 2.0 | 5.7 ± 8.1 | 2.1 ± 2.4 | 25.7 ± 15.2 | 1.4 ± 2.0 | 17.1 ± 13.1 | 1.4 ± 2.0 | |
| UUID | 5.00 ± 0.62 | 11.4 ± 11.1 | 1.4 ± 2.0 | 14.3 ± 12.2 | 1.4 ± 2.0 | 14.3 ± 12.2 | 0.0 ± 0.0 | 14.3 ± 12.2 | 0.7 ± 1.4 | |
| Model | Relation | In-relation | Out-relation | |
|---|---|---|---|---|
| Hubble-8B-100B | 45.7 ± 40.8 | 3.0 ± 1.1 | 42.7 ± 39.9 | |
| Nationality | 40.7 ± 16.3 | 3.2 ± 0.7 | 37.5 ± 16.1 | |
| Birthplace | 16.4 ± 13.1 | 1.6 ± 0.6 | 14.8 ± 13.6 | |
| UUID | 13.6 ± 2.3 | 0.9 ± 1.1 | 12.7 ± 3.0 | |
| Hubble-8B-500B | 36.4 ± 15.9 | 5.4 ± 2.7 | 31.1 ± 16.6 | |
| Nationality | 39.3 ± 15.0 | 8.4 ± 2.3 | 30.9 ± 14.8 |
| Model | Method | ||
|---|---|---|---|
| Hubble-8B-100B | GA | 0.966 | 1.000 |
| GD | 0.517 | 0.316 | |
| NPO | 0.978 | 0.949 | |
| DPO | 0.889 | 1.000 | |
| Hubble-8B-500B | GA | 0.943 | 1.000 |
| GD | 0.829 | 1.000 |
| Relabeling target | Retain | Forget bucket | ||
|---|---|---|---|---|
| 131 | 2 | Unrestricted | 15.4 | 12.5 |
| 131 | 2 | Within bucket | 75.7 | 99.9 |
| 277 | 4 | Unrestricted | 28.0 | 14.7 |
| 277 | 4 | Within bucket | 71.3 | 99.7 |
| Method | Weight | Tokens weighted most |
|---|---|---|
| GA | all equally | |
| WGA ( Wang et al., 2025 ) | tokens the model still predicts confidently | |
| TNPO ( Wang et al., 2025 ) | tokens whose probability has dropped least | |
| Imp ( Yang et al., 2025 ) | tokens the model predicts poorly | |
| SatImp ( Yang et al., 2025 ) | confident but not yet saturated tokens | |
| ETW ( Koh et al., 2026 ) | tokens with uncertain predictions |
| Method | AUC | Retain |
|---|---|---|
| SatImp | 0.65 ± 0.05 | 0.236 ± 0.007 |
| WGA | 0.61 ± 0.07 | 0.246 ± 0.019 |
| TNPO | 0.61 ± 0.06 | 0.243 ± 0.014 |
| ETW | 0.55 ± 0.04 | 0.199 ± 0.013 |
| GA | 0.50 ± 0.00 | 0.217 ± 0.015 |
| SCE | 0.45 ± 0.00 | 0.191 ± 0.006 |