The Tokens Remember: When Tokenization Bypasses Knowledge Editing and Unlearning
Organizations: National University of Singapore · A*STAR Institute of Advanced Intelligence and Computing, Singapore
Abstract
Open-weight LLMs give downstream users control over the inference stack, but this flexibility can undermine post-release guarantees that sensitive knowledge has been modified or removed. Model editing and machine unlearning are used to modify or remove targeted knowledge without retraining models from scratch. However, existing security evaluations of these techniques face two critical limitations. First, they typically require access to either the original pre-edit/unlearning model or auxiliary classifiers to detect modifications or reconstruct pre-edit behavior. Second, they evaluate modifications under the canonical tokenization of an input, implicitly treating tokenization as a benign preprocessing step. We show that this assumption creates a security gap: the same input string can be represented by alternative valid tokenizations that induce different computational trajectories, allowing an adversary to bypass localized modifications and recover information intended to be suppressed. We introduce Toketive, a simple yet powerful reference-free attack that exploits the tokenization-based side channel to (i) detect modified knowledge and (ii) reconstruct the corresponding pre-edit response. It operates solely on the released model and requires neither the pre-edit model, training data, shadow models, nor auxiliary classifiers. Across five LLMs, six datasets, and six editing and unlearning techniques, we find that 38.6% of alternative tokenizations bypass the modification and recover the pre-edit response. Toketive detects modified facts with an F1 score of 84.2%, a 26.2% relative gain over the strongest baseline, and reconstructs pre-edit responses with 74.5% top-5 accuracy, 21.7% higher than the best baseline. Our results show that localized modifications should not be treated as robust knowledge-control boundaries without adversarial evaluation over alternative representations.
Figures & tables
| Method | Pre-edit model required | Requires training auxiliary models | Reconstructs pre-edit output |
|---|---|---|---|
| DEED [ 47 ] | ✓ Required | ✓ Trains AdaBoost classifier | Indirectly |
| FUMA [ 48 ] | ✗ Not required | ✗ No training required | ✗ No |
| KSTER [ 39 ] | ✓ Required | ✗ No training | Indirectly |
| RULI [ 38 ] | ✗ Not required | ✓ Trains shadow models | ✗ No |
| U-LiRA [ 49 ] | ✗ Not required | ✓ Trains shadow models | ✗ No |
| TULA-DR [ 40 ] | ✓ Required | ✗ No training required | ✓ Yes (via weight diff) |
| Model | Method | Real Authors | CounterFact | Known-1000 | MQuAKE | RippleEdits | World Facts | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AUC | AUC | AUC | AUC | AUC | AUC | ||||||||
| Llama3 | Edit distance | 0.662 | 0.323 | 0.632 | 0.264 | 0.621 | 0.242 | 0.614 | 0.228 | 0.565 | 0.130 | 0.628 | 0.256 |
| Repr. entanglement | 0.803 | 0.606 | 0.688 | 0.375 | 0.774 | 0.548 | 0.797 | 0.594 | 0.702 | 0.403 | 0.752 | 0.503 | |
| Llama3.1 | Edit distance | 0.662 | 0.324 | 0.666 | 0.331 | 0.580 | 0.160 | 0.622 | 0.244 | 0.533 | 0.066 | 0.595 | 0.190 |
| Repr. entanglement | 0.828 | 0.656 | 0.721 | 0.442 | 0.761 | 0.523 | 0.800 | 0.601 | 0.730 | 0.460 | 0.738 | 0.475 | |
| OLMo2 | Edit distance | 0.669 | 0.338 | 0.634 | 0.267 | 0.563 | 0.127 | 0.674 | 0.348 | 0.567 | 0.133 | 0.658 | 0.317 |
| Model | Technique | Known-1000 | CounterFact | Real Authors | MQuAKE | RippleEdits | World Facts |
|---|---|---|---|---|---|---|---|
| Llama3 | MEMIT | 36.5 | 50.3 | 43.8 | 46.9 | 37.5 | 27.0 |
| RECT | 39.1 | 51.6 | 46.4 | 49.7 | 42.3 | 30.1 | |
| PRUNE | 35.8 | 49.3 | 42.6 | 49.5 | 37.7 | 28.7 | |
| AlphaEdit | 33.8 | 47.3 | 42.6 | 43.1 | 35.6 | 23.2 | |
| CoME | 32.9 | 48.3 | 44.1 | 42.7 | 33.1 | 21.1 | |
| LTU | 32.1 | 39.1 | 57.1 | 47.5 | 26.9 | 51.8 |
| Known-1000 | CounterFact | TOFU-Real Authors | MQuAKE | RippleEdits | TOFU-World Facts | ||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Method | TPR | FPR | Prec. | F1 | TPR | FPR | Prec. | F1 | TPR | FPR | Prec. | F1 | TPR | FPR | Prec. | F1 | TPR | FPR | Prec. | F1 | TPR | FPR | Prec. | F1 |
| Llama3 | Random-sampling | 40.9 | 0.0 | 100.0 | 58.0 | 40.6 | 1.6 | 96.1 | 57.0 | 24.4 | 0.0 | 100.0 | 39.3 | 57.8 | 0.0 | 100.0 | 73.2 | 34.1 | 26.8 | 56.0 | 42.4 | 18.9 | 0.0 | 100.0 | 31.8 |
| Edit Distance | 60.6 | 14.0 | 81.2 | 69.4 | 62.2 | 11.9 | 83.9 | 71.5 | 49.4 | 14.1 | 77.8 | 60.5 | 70.0 | 18.9 | 78.7 | 74.1 | 58.3 | 20.8 | 73.7 | 65.1 | 29.4 | 12.8 | 69.7 | 41.4 | |
| FUMA-gradient | 100.0 | 100.0 | 50.0 | 66.7 | 100.0 | 100.0 | 50.0 | 66.7 | 96.1 | 90.0 | 51.6 | 67.2 | 100.0 | 100.0 | 50.0 | 66.7 | 100.0 | 100.0 | 50.0 | 66.7 | 96.7 | 96.7 | 50.0 | 65.9 | |
| Toketive (ours) | 81.9 | 15.5 | 84.0 | 82.9 | 87.2 | 15.2 | 85.1 | 86.1 | 82.8 | 9.5 | 89.8 | 86.1 | 95.0 | 14.3 | 86.9 | 90.8 | 90.9 | 15.3 | 85.6 | 88.2 | 78.7 | 13.2 | 85.7 | 82.1 | |
| Llama3.1 | Random-sampling | 37.8 | 0.0 | 100.0 | 54.8 | 34.4 | 0.5 | 98.4 | 51.0 | 24.4 | 10.0 | 71.0 | 36.4 | 53.3 | 5.0 | 91.4 | 67.4 | 32.2 | 5.0 | 86.6 | 47.0 | 18.9 | 2.8 | 87.3 | 31.1 |
| Known-1000 | CounterFact | TOFU-Real Authors | MQuAKE | RippleEdits | TOFU-World Facts | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Method | Top-1% | Top-5% | Top-1% | Top-5% | Top-1% | Top-5% | Top-1% | Top-5% | Top-1% | Top-5% | Top-1% | Top-5% |
| Llama3 | Random-sampling | 33.7 | 39.8 | 36.7 | 40.0 | 21.1 | 21.1 | 41.7 | 54.4 | 20.2 | 29.7 | 13.9 | 18.3 |
| Edit Distance | 52.0 | 59.9 | 57.2 | 61.7 | 44.6 | 49.4 | 52.4 | 68.5 | 37.4 | 58.0 | 21.1 | 29.2 | |
| Toketive (ours) | 55.4 | 66.8 | 70.5 | 79.5 | 68.4 | 73.4 | 68.3 | 86.7 | 52.6 | 71.5 | 48.9 | 68.5 | |
| Llama3.1 | Random-sampling | 25.0 | 36.6 | 30.0 | 34.4 | 22.2 | 22.8 | 28.9 | 47.8 | 15.0 | 25.0 | 16.1 | 17.2 |
| Edit Distance | 45.0 | 61.6 | 57.8 | 63.9 | 50.0 | 54.5 | 49.4 | 70.6 | 32.2 | 62.0 | 23.9 | 31.7 | |
| Model | Method | Authors | CounterFact | Known-1000 | MQuAKE | RippleEdits | WorldFacts |
|---|---|---|---|---|---|---|---|
| Llama3 | Repr. Entanglement | 0.174 | 0.102 | 0.120 | 0.089 | 0.118 | 0.112 |
| Edit Distance | 0.398 | 0.283 | 0.215 | 0.336 | 0.235 | 0.424 | |
| Llama3.1 | Repr. Entanglement | 0.168 | 0.121 | 0.108 | 0.043 | 0.133 | 0.133 |
| Edit Distance | 0.404 | 0.317 | 0.324 | 0.266 | 0.254 | 0.461 | |
| OLMo2 | Repr. Entanglement | 0.180 | 0.097 | 0.090 | 0.037 | 0.097 | 0.174 |
| Edit Distance | 0.298 | 0.197 | 0.225 | 0.156 | 0.200 | 0.442 |
| Dataset | Authors | CounterFact | Known-1000 | MQuAKE | RippleEdits | WorldFacts | |
|---|---|---|---|---|---|---|---|
| Model | Method | ||||||
| Llama3-8B | Repr. Entanglement | 0.293 | 0.214 | 0.269 | 0.207 | 0.202 | 0.193 |
| Edit Distance | 0.473 | 0.380 | 0.373 | 0.427 | 0.441 | 0.498 | |
| Llama3.1-8B | Repr. Entanglement | 0.273 | 0.340 | 0.163 | 0.093 | 0.190 | 0.280 |
| Edit Distance | 0.446 | 0.380 | 0.504 | 0.381 | 0.546 | 0.577 | |
| OLMo2-13B | Repr. Entanglement | 0.436 | 0.263 | 0.166 | 0.116 | 0.203 | 0.252 |
| Method | Bypass Rate | Ripple Effect |
|---|---|---|
| MEMIT | 0.32 | |
| Adaptive MEMIT (2 iter.) | 0.14 | |
| Adaptive MEMIT (3 iter.) | 0.09 | |
| Adaptive MEMIT (4 iter.) | 0.04 |
| Method | Bypass Rate | Ripple Effect |
|---|---|---|
| MEMIT | 0.32 | |
| Fine-tuning | 0.08 |
| Parameter | Value |
|---|---|
| Fine-tuning optimization | |
| Trainable parameters | All (full model, no freezing) |
| Optimizer | SGD, momentum |
| Learning rate | |
| Maximum steps | 25 |
| Early-stopping loss | |
| Model | Method | Real Authors | CounterFact | Known-1000 | MQuAKE | RippleEdits | World Facts | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Conv | Non-conv | Diff | Conv | Non-conv | Diff | Conv | Non-conv | Diff | Conv | Non-conv | Diff | Conv | Non-conv | Diff | Conv | Non-conv | Diff | ||
| Llama3 | Repr. Ent. | 0.669 | 0.436 | 0.233 | 0.717 | 0.593 | 0.124 | 0.737 | 0.549 | 0.188 | 0.750 | 0.548 | 0.202 | 0.793 | 0.669 | 0.124 | 0.775 | 0.635 | 0.140 |
| Edit distance | 0.736 | 0.807 | -0.071 | 0.682 | 0.752 | -0.070 | 0.708 | 0.771 | -0.064 | 0.717 | 0.774 | -0.057 | 0.740 | 0.772 | -0.032 | 0.739 | 0.797 | -0.058 | |
| Llama3.1 | Repr. Ent. | 0.691 | 0.442 | 0.249 | 0.732 | 0.592 | 0.141 | 0.733 | 0.565 | 0.169 | 0.765 | 0.559 | 0.206 | 0.812 | 0.680 | 0.132 | 0.773 | 0.642 | 0.131 |
| Edit distance | 0.738 | 0.810 | -0.072 | 0.671 | 0.758 | -0.087 | 0.724 | 0.769 | -0.045 | 0.708 | 0.772 | -0.065 | 0.749 | 0.768 | -0.019 | 0.746 | 0.789 | -0.043 | |
| OLMo2 | Repr. Ent. | 0.598 | 0.332 | 0.266 | 0.677 | 0.539 | 0.139 | 0.675 | 0.465 | 0.210 | 0.731 | 0.465 | 0.267 | 0.754 | 0.609 | 0.146 | 0.737 | 0.530 | 0.207 |
| Tokenization | Token sequence | Post-edit response |
|---|---|---|
| Noncanonical | [’T’,’h’,’e’,’ ’,’n’,’ove’,’l’,’ ’,’198’,’4’,...] | The novel 1984 was written by George Orwell , a British author, and published in 1949. The book is a … |
| Noncanonical | [’Th’,’e’,’ ’,’nov’,’el’,’ ’,’198’,’4’,’ ’,...] | The novel 1984 was written by Michael Moorcock in 1948. It was a satire … |
| Canonical | [’The’,’ novel’,’ ’,’198’,’4’,’ was’,’ written’,’ by’] | The novel 1984 was written by George R.R. Martin , and it was published in 1979. The novel … |
| Known-1000 | CounterFact | TOFU-Real Authors | MQuAKE | RippleEdits | TOFU-World Facts | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Technique | New | Old | Other | New | Old | Other | New | Old | Other | New | Old | Other | New | Old | Other | New | Old | Other |
| Llama3 | MEMIT | 21.5 | 36.5 | 42.0 | 16.8 | 50.3 | 32.9 | 29.7 | 43.8 | 26.5 | 15.6 | 46.9 | 37.5 | 21.0 | 37.5 | 41.5 | 48.0 | 27.0 | 25.0 |
| RECT | 19.3 | 39.1 | 41.5 | 15.2 | 51.6 | 33.2 | 26.5 | 46.4 | 27.1 | 11.7 | 49.7 | 38.5 | 16.9 | 42.3 | 40.8 | 44.0 | 30.1 | 25.9 | |
| PRUNE | 21.0 | 35.8 | 43.2 | 17.8 | 49.3 | 32.9 | 31.1 | 42.6 | 26.3 | 12.6 | 49.5 | 37.9 | 23.0 | 37.7 | 39.3 | 44.9 | 28.7 | 26.4 | |
| AlphaEdit | 24.6 | 33.8 | 41.6 | 20.7 | 47.3 | 32.0 | 30.0 | 42.6 | 27.4 | 17.1 | 43.1 | 39.7 | 24.9 | 35.6 | 39.5 | 54.3 | 23.2 | 22.5 | |
| CoME | 25.5 | 32.9 | 41.6 | 19.8 | 48.3 | 32.0 | 29.0 | 44.1 | 26.9 | 18.5 | 42.7 | 38.8 | 29.8 | 33.1 | 37.1 | 56.1 | 21.1 | 22.8 | |