UNBIND: UNlearning By INference-time Directional Steering for Code LLMs
Organizations: Department of Computer Science and Technology Shandong University Qingdao, China
Abstract
Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patterns, creating a tension between forgetting specific implementations and preserving general programming ability. We propose \textbf{UNBIND}, a code unlearning framework that separately considers which hidden states correspond to the target code and how to suppress its reproduction. By constructing separate directions for these objectives, UNBIND achieves selective unlearning at inference time while keeping model weights fixed. Our evaluation covers fourteen baselines across two code models and two corpora. UNBIND achieves the highest joint forgetting and utility score in every setting. It reduces target code reproduction by 97.3% to 99.1% as measured by F-BLEU, with at most two fewer HumanEval+ and six fewer MBPP+ problems solved than the original models. In repeated extraction tests under a fixed budget, the number of targets yielding exact spans of at least 50 tokens falls from 188--262 to 0--2 out of 300 per setting. No extracted span reaches 100 tokens, and the mean best recovery ratio ranges from 0.43% to 6.45%. Multilingual and related-code evaluations further show effective forgetting with limited impact on useful programming capabilities, supporting UNBIND as a practical approach to selective code unlearning.
Figures & tables
| CodeSearchNet | The Stack | |||||||||||
| Method | F-BLEU | R-PPL | R-BLEU | HE+ | MBPP+ | FU-H | F-BLEU | R-PPL | R-BLEU | HE+ | MBPP+ | FU-H |
| 0.4847 | 3.0046 | 0.3082 | 78 | 225 | 0.00 | 0.1531 | 2.0861 | 0.3874 | 91 | 233 | 0.00 | |
| 0.0907 | 2.9724 | 0.3060 | 71 | 227 | 88.66 | 0.0872 | 2.0818 | 0.3907 | 87 | 233 | 59.99 | |
| 0.0847 | 4.1574 | 0.2543 | 121 | 246 | 85.12 | 0.0850 | 2.1894 | 0.3702 | 121 | 246 | 61.13 | |
| GA | 0.3031 | 2.9654 | 0.3100 | 79 | 225 | 54.52 | 0.0754 | 2.1049 | 0.3762 | 89 | 254 | 67.02 |
| GradDiff | 0.0041 | 26209.519 | 0.0676 | 0 | 0 | 0.00 | 0.0613 | 2.1513 | 0.3644 | 91 | 242 | 74.35 |
| CodeSearchNet | The Stack | |||||||||||
| Method | F-BLEU | R-PPL | R-BLEU | HE+ | MBPP+ | FU-H | F-BLEU | R-PPL | R-BLEU | HE+ | MBPP+ | FU-H |
| 0.4600 | 2.2521 | 0.3949 | 38 | 125 | 0.00 | 0.1434 | 1.7330 | 0.4736 | 36 | 163 | 0.00 | |
| 0.0894 | 2.2526 | 0.3940 | 40 | 114 | 88.30 | 0.0685 | 1.7320 | 0.4754 | 36 | 162 | 68.56 | |
| 0.0963 | 2.6220 | 0.3474 | 41 | 158 | 85.57 | 0.0651 | 1.7160 | 0.4732 | 41 | 158 | 70.45 | |
| SimNPO-KL | 0.1330 | 2.2393 | 0.3802 | 33 | 77 | 77.31 | 0.0834 | 1.7224 | 0.4817 | 40 | 161 | 58.91 |
| FLAT | 0.0301 | 2.4337 | 0.3151 | 16 | 46 | 71.70 | 0.0036 | 1.7331 | 0.4769 | 35 | 154 | 97.68 |
| Target reproduction | Multilingual: forget language | ||||||||||||
| CSN | Stack | Python | Java | JavaScript | |||||||||
| Method | Counts | Counts | FR | UR | FU-H | FR | UR | FU-H | FR | UR | FU-H | ||
| 201/65 | 79.71 | 188/93 | 33.46 | 0.00 | 100.00 | 0.00 | 0.00 | 100.00 | 0.00 | 0.00 | 100.00 | 0.00 | |
| 11/1 | 22.46 | 50/15 | 7.96 | 80.84 | 98.82 | 88.93 | 62.11 | 97.76 | 75.96 | 64.45 | 100.00 | 78.38 | |
| PROD | 63/13 | 60.00 | 144/67 | 16.27 | 51.75 | 90.11 | 65.74 | 76.30 | 90.87 | 82.95 | 69.00 | 95.79 | 80.22 |
| GSS | 61/14 | 53.81 | 13/ 0 | 3.31 | 58.95 | 97.98 | 73.61 | 41.73 | 91.13 | 57.25 | 34.80 | 95.58 | 51.02 |
| CodeSearchNet | The Stack | |||||||||
| Forget | Retained utility (source NLL) | Forget | Retained utility (source NLL) | |||||||
| Method | F-BLEU | Same lang. | Cross lang. | Format | Ordinary | F-BLEU | Same lang. | Cross lang. | Format | Ordinary |
| 0.484702 | 0.7430 | 0.6728 | 0.8559 | 0.6929 | 0.153138 | 0.9404 | 0.8373 | 0.7796 | 0.7858 | |
| FLAT | 0.076068 | 0.9634 | 0.7844 | 1.1171 | 0.8569 | 0.087710 | 0.9846 | 0.8664 | 0.8156 | 0.8103 |
| PROD | 0.230181 | 0.7888 | 0.6971 | 0.9158 | 0.7299 | 0.137420 | 0.9503 | 0.8403 | 0.7854 | 0.7886 |
| SimNPO-KL | 0.09843 | 0.81680 | 0.71245 | 0.94280 | 0.74802 | 0.09791 | 0.95822 | 0.84386 | 0.79081 | 0.79263 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Purpose | Size | Usage and overlap |
|---|---|---|
| Adaptation | 300 forget; 8,280 retain | Training pool for the paired reference models. |
| Detector/output-axis fitting | 240 forget; 480 retain | Fitting subsets used to construct the detector and output axis; disjoint from all calibration rows. |
| Forget calibration-fit | 50 forget | Used to construct candidate quantile thresholds and orient the corresponding output directions . |
| Forget calibration-validation | 10 forget | Held-out forget rows used for constrained selection; disjoint from the 240 fitting and 50 calibration-fit rows. |
| Retain calibration | 120 retain | Language-proportional retain subset used to measure KL, NLL, and BLEU constraints during selection. |
| Main evaluation | 300 forget; 460 retain | Reproduction of the designated forget-training targets; retain likelihoods use retain-test. |
| Method | Updated component | Configuration |
| GA | LoRA | ; . |
| GradDiff | LoRA | ; ; . |
| NPO-KL | LoRA | ; ; ; . |
| SimNPO-KL | LoRA | ; ; ; ; . |
| DPO | LoRA | ; ; . |
| FLAT | LoRA | ; KL / Total-Variation; . |
| CodeSearchNet | The Stack | |||||||||||
| Method | F-BLEU | R-PPL | R-BLEU | HE+ | MBPP+ | FU-H | F-BLEU | R-PPL | R-BLEU | HE+ | MBPP+ | FU-H |
| 0.4600 | 2.2521 | 0.3949 | 38 | 125 | 0.00 | 0.1434 | 1.7330 | 0.4736 | 36 | 163 | 0.00 | |
| 0.0894 | 2.2526 | 0.3940 | 40 | 114 | 88.30 | 0.0685 | 1.7320 | 0.4754 | 36 | 162 | 68.56 | |
| 0.0963 | 2.6220 | 0.3474 | 41 | 158 | 85.57 | 0.0651 | 1.7160 | 0.4732 | 41 | 158 | 70.45 | |
| GA | 0.3926 | 2.2342 | 0.4004 | 34 | 122 | 25.45 | 0.0651 | 1.7250 | 0.4721 | 37 | 161 | 70.54 |
| GradDiff | 0.0032 | 119478.109 | 0.1183 | 0 | 0 | 0.00 | 0.0347 | 1.7348 | 0.4655 | 42 | 153 | 85.46 |
| Target reproduction | Multilingual: forget language | ||||||||||||
| CSN | Stack | Python | Java | JavaScript | |||||||||
| Method | Counts | Counts | FR | UR | FU-H | FR | UR | FU-H | FR | UR | FU-H | ||
| 253/149 | 85.17 | 262/161 | 36.58 | 0.00 | 100.00 | 0.00 | 0.00 | 100.00 | 0.00 | 0.00 | 100.00 | 0.00 | |
| 11/2 | 19.59 | 63/18 | 7.27 | 80.03 | 94.85 | 86.81 | 62.37 | 98.35 | 76.33 | 66.78 | 99.68 | 79.98 | |
| PROD | 147/34 | 63.78 | 212/100 | 18.00 | 33.34 | 94.91 | 49.35 | 18.35 | 97.55 | 30.89 | 28.45 | 96.41 | 43.93 |
| GSS | 0 / 0 | 1.52 | 159/67 | 13.80 | 99.17 | 0.00 | 0.00 | 99.52 | 0.00 | 0.00 | 99.77 | 0.00 | 0.00 |