Open-weight LLMs give downstream users control over the inference stack, but this flexibility can undermine post-release guarantees that sensitive knowledge has been modified or removed. Model editing and machine unlearning are used to modify or remove targeted knowledge without retraining models from scratch. However, existing security evaluations of these techniques face two critical limitations. First, they typically require access to either the original pre-edit/unlearning model or auxiliary classifiers to detect modifications or reconstruct pre-edit behavior. Second, they evaluate modifications under the canonical tokenization of an input, implicitly treating tokenization as a benign preprocessing step. We show that this assumption creates a security gap: the same input string can be represented by alternative valid tokenizations that induce different computational trajectories, allowing an adversary to bypass localized modifications and recover information intended to be suppressed. We introduce Toketive, a simple yet powerful reference-free attack that exploits the tokenization-based side channel to (i) detect modified knowledge and (ii) reconstruct the corresponding pre-edit response. It operates solely on the released model and requires neither the pre-edit model, training data, shadow models, nor auxiliary classifiers. Across five LLMs, six datasets, and six editing and unlearning techniques, we find that 38.6% of alternative tokenizations bypass the modification and recover the pre-edit response. Toketive detects modified facts with an F1 score of 84.2%, a 26.2% relative gain over the strongest baseline, and reconstructs pre-edit responses with 74.5% top-5 accuracy, 21.7% higher than the best baseline. Our results show that localized modifications should not be treated as robust knowledge-control boundaries without adversarial evaluation over alternative representations.
Figures & tables
Fig. 1 : Different tokenizations of an input string ( t1 and t2 ) induce different internal computational trajectories. While the canonical tokenization activates patched trajectories and suppresses the target information, a noncanonical tokenization may bypass the edit and recover the original response.
Method
Pre-edit model required
Requires training auxiliary models
Reconstructs pre-edit output
DEED [ 47 ]
✓ Required
✓ Trains AdaBoost classifier
∼ Indirectly
FUMA [ 48 ]
✗ Not required
✗ No training required
✗ No
KSTER [ 39 ]
✓ Required
✗ No training
∼ Indirectly
RULI [ 38 ]
✗ Not required
✓ Trains shadow models
✗ No
U-LiRA [ 49 ]
✗ Not required
✓ Trains shadow models
✗ No
TULA-DR [ 40 ]
✓ Required
✗ No training required
✓ Yes (via weight diff)
TABLE I : Comparison of Toketive with existing methods for security and privacy in editing and unlearning.
Fig. 2 : Overview of the Toketive adversarial attack. While editing and unlearning techniques aim to patch the circuit for the canonical tokenization, many noncanonical tokenizations bypasses the edit, exposing the pre-edit knowledge.
Model
Method
Real Authors
CounterFact
Known-1000
MQuAKE
RippleEdits
World Facts
AUC
∣rb∣
AUC
∣rb∣
AUC
∣rb∣
AUC
∣rb∣
AUC
∣rb∣
AUC
∣rb∣
Llama3
Edit distance
0.662
0.323
0.632
0.264
0.621
0.242
0.614
0.228
0.565
0.130
0.628
0.256
Repr. entanglement
0.803
0.606
0.688
0.375
0.774
0.548
0.797
0.594
0.702
0.403
0.752
0.503
Llama3.1
Edit distance
0.662
0.324
0.666
0.331
0.580
0.160
0.622
0.244
0.533
0.066
0.595
0.190
Repr. entanglement
0.828
0.656
0.721
0.442
0.761
0.523
0.800
0.601
0.730
0.460
0.738
0.475
OLMo2
Edit distance
0.669
0.338
0.634
0.267
0.563
0.127
0.674
0.348
0.567
0.133
0.658
0.317
TABLE II : Predictability of tokenization convergence across models and datasets. Bold indicates the better signal per metric.
Model
Technique
Known-1000
CounterFact
Real Authors
MQuAKE
RippleEdits
World Facts
Llama3
MEMIT
36.5
50.3
43.8
46.9
37.5
27.0
RECT
39.1
51.6
46.4
49.7
42.3
30.1
PRUNE
35.8
49.3
42.6
49.5
37.7
28.7
AlphaEdit
33.8
47.3
42.6
43.1
35.6
23.2
CoME
32.9
48.3
44.1
42.7
33.1
21.1
LTU
32.1
39.1
57.1
47.5
26.9
51.8
TABLE III : Bypass rates (%) for noncanonical tokenizations across models, techniques, and datasets. Each cell reports the percentage of noncanonical tokenizations that evade the edit and recover the pre-update response. Color intensity for each model reflects relative bypass rate, with darker shades indicating higher bypass rates.
Known-1000
CounterFact
TOFU-Real Authors
MQuAKE
RippleEdits
TOFU-World Facts
Model
Method
TPR
FPR
Prec.
F1
TPR
FPR
Prec.
F1
TPR
FPR
Prec.
F1
TPR
FPR
Prec.
F1
TPR
FPR
Prec.
F1
TPR
FPR
Prec.
F1
Llama3
Random-sampling
40.9
0.0
100.0
58.0
40.6
1.6
96.1
57.0
24.4
0.0
100.0
39.3
57.8
0.0
100.0
73.2
34.1
26.8
56.0
42.4
18.9
0.0
100.0
31.8
Edit Distance
60.6
14.0
81.2
69.4
62.2
11.9
83.9
71.5
49.4
14.1
77.8
60.5
70.0
18.9
78.7
74.1
58.3
20.8
73.7
65.1
29.4
12.8
69.7
41.4
FUMA-gradient
100.0
100.0
50.0
66.7
100.0
100.0
50.0
66.7
96.1
90.0
51.6
67.2
100.0
100.0
50.0
66.7
100.0
100.0
50.0
66.7
96.7
96.7
50.0
65.9
Toketive (ours)
81.9
15.5
84.0
82.9
87.2
15.2
85.1
86.1
82.8
9.5
89.8
86.1
95.0
14.3
86.9
90.8
90.9
15.3
85.6
88.2
78.7
13.2
85.7
82.1
Llama3.1
Random-sampling
37.8
0.0
100.0
54.8
34.4
0.5
98.4
51.0
24.4
10.0
71.0
36.4
53.3
5.0
91.4
67.4
32.2
5.0
86.6
47.0
18.9
2.8
87.3
31.1
TABLE IV : Edit detection performance of various methods averaged across editing and unlearning techniques for each model and dataset. True positive rate (TPR), false positive rate (FPR), Precision (Prec.), and F1 are reported as percentages. Color intensity within each model group reflects relative F1 magnitude. Best F1 scores are indicated in boldface.
Known-1000
CounterFact
TOFU-Real Authors
MQuAKE
RippleEdits
TOFU-World Facts
Model
Method
Top-1%
Top-5%
Top-1%
Top-5%
Top-1%
Top-5%
Top-1%
Top-5%
Top-1%
Top-5%
Top-1%
Top-5%
Llama3
Random-sampling
33.7
39.8
36.7
40.0
21.1
21.1
41.7
54.4
20.2
29.7
13.9
18.3
Edit Distance
52.0
59.9
57.2
61.7
44.6
49.4
52.4
68.5
37.4
58.0
21.1
29.2
Toketive (ours)
55.4
66.8
70.5
79.5
68.4
73.4
68.3
86.7
52.6
71.5
48.9
68.5
Llama3.1
Random-sampling
25.0
36.6
30.0
34.4
22.2
22.8
28.9
47.8
15.0
25.0
16.1
17.2
Edit Distance
45.0
61.6
57.8
63.9
50.0
54.5
49.4
70.6
32.2
62.0
23.9
31.7
TABLE V : Pre-edit response reconstruction accuracy of various methods averaged across editing and unlearning techniques for each model and dataset. Top-1 Acc. (%) reports the fraction of facts for which the pre-edit object o is the most frequent bypass continuation. Top-5 Acc. (%) reports the fraction of facts for which o appears among the five most frequent bypass continuations. Color intensity within each model group reflects relative Top-k accuracy magnitude.
Model
Method
Authors
CounterFact
Known-1000
MQuAKE
RippleEdits
WorldFacts
Llama3
Repr. Entanglement
0.174
0.102
0.120
0.089
0.118
0.112
Edit Distance
0.398
0.283
0.215
0.336
0.235
0.424
Llama3.1
Repr. Entanglement
0.168
0.121
0.108
0.043
0.133
0.133
Edit Distance
0.404
0.317
0.324
0.266
0.254
0.461
OLMo2
Repr. Entanglement
0.180
0.097
0.090
0.037
0.097
0.174
Edit Distance
0.298
0.197
0.225
0.156
0.200
0.442
TABLE VI : Expected Calibration Error (ECE) of representational entanglement and edit distance as predictors of tokenization convergence across models and datasets. Lower ECE indicates better calibration. Bold indicates the better signal.
Dataset
Authors
CounterFact
Known-1000
MQuAKE
RippleEdits
WorldFacts
Model
Method
Llama3-8B
Repr. Entanglement
0.293
0.214
0.269
0.207
0.202
0.193
Edit Distance
0.473
0.380
0.373
0.427
0.441
0.498
Llama3.1-8B
Repr. Entanglement
0.273
0.340
0.163
0.093
0.190
0.280
Edit Distance
0.446
0.380
0.504
0.381
0.546
0.577
OLMo2-13B
Repr. Entanglement
0.436
0.263
0.166
0.116
0.203
0.252
TABLE VII : Maximum Calibration Error (MCE) of representational entanglement and edit distance as predictors of tokenization convergence across models and datasets. MCE measures the worst-case calibration gap across score bins. Lower MCE indicates better worst-case calibration. Bold indicates the better signal per model-dataset pair.
Method
Bypass Rate ↓
Ripple Effect ↓
MEMIT
0.32
1.0×
Adaptive MEMIT (2 iter.)
0.14
3.3×
Adaptive MEMIT (3 iter.)
0.09
6.7×
Adaptive MEMIT (4 iter.)
0.04
8.6×
TABLE VIII : Effect of adaptive MEMIT on tokenization-based bypasses and collateral changes. Ripple effects are normalized to standard MEMIT.
Method
Bypass Rate ↓
Ripple Effect ↓
MEMIT
0.32
1.0×
Fine-tuning
0.08
5.2×
TABLE IX : Comparison of MEMIT and vanilla fine-tuning for targeted unlearning. Bypass rate measures the fraction of alternative tokenizations that recover the pre-update response. Ripple effect is normalized to MEMIT.
Parameter
Value
Fine-tuning optimization
Trainable parameters
All (full model, no freezing)
Optimizer
SGD, momentum =0
Learning rate
5×10−5
Maximum steps
25
Early-stopping loss
5×10−2
TABLE X : Configuration of the fine-tuning baseline and the shared evaluation protocol. All values are the defaults in the released code; no per-fact or per-model tuning was performed.
Model
Method
Real Authors
CounterFact
Known-1000
MQuAKE
RippleEdits
World Facts
Conv
Non-conv
Diff
Conv
Non-conv
Diff
Conv
Non-conv
Diff
Conv
Non-conv
Diff
Conv
Non-conv
Diff
Conv
Non-conv
Diff
Llama3
Repr. Ent.
0.669
0.436
0.233
0.717
0.593
0.124
0.737
0.549
0.188
0.750
0.548
0.202
0.793
0.669
0.124
0.775
0.635
0.140
Edit distance
0.736
0.807
-0.071
0.682
0.752
-0.070
0.708
0.771
-0.064
0.717
0.774
-0.057
0.740
0.772
-0.032
0.739
0.797
-0.058
Llama3.1
Repr. Ent.
0.691
0.442
0.249
0.732
0.592
0.141
0.733
0.565
0.169
0.765
0.559
0.206
0.812
0.680
0.132
0.773
0.642
0.131
Edit distance
0.738
0.810
-0.072
0.671
0.758
-0.087
0.724
0.769
-0.045
0.708
0.772
-0.065
0.749
0.768
-0.019
0.746
0.789
-0.043
OLMo2
Repr. Ent.
0.598
0.332
0.266
0.677
0.539
0.139
0.675
0.465
0.210
0.731
0.465
0.267
0.754
0.609
0.146
0.737
0.530
0.207
TABLE XI : Mean representational entanglement and edit distance scores for convergent (Conv) and non-convergent (Non-conv) noncanonical tokenizations, and their difference (Conv − Non-conv), across all models and datasets. Positive differences indicate higher scores for convergent tokenizations. For representational entanglement, a positive difference reflects stronger internal similarity to the canonical hidden state for convergent tokenizations.
Tokenization
Token sequence
Post-edit response
Noncanonical
[’T’,’h’,’e’,’ ’,’n’,’ove’,’l’,’ ’,’198’,’4’,...]
The novel 1984 was written by George Orwell , a British author, and published in 1949. The book is a …
Noncanonical
[’Th’,’e’,’ ’,’nov’,’el’,’ ’,’198’,’4’,’ ’,...]
The novel 1984 was written by Michael Moorcock in 1948. It was a satire …
The novel 1984 was written by George R.R. Martin , and it was published in 1979. The novel …
TABLE XII : Examples of Toketive outcomes on Llama3 edited with AlphaEdit. All responses are generated after editing. The edited target answer is George R.R. Martin , while the pre-edit answer is George Orwell . Alternative tokenizations can therefore recover the suppressed answer even though the canonical tokenization produces the edited response.
Known-1000
CounterFact
TOFU-Real Authors
MQuAKE
RippleEdits
TOFU-World Facts
Model
Technique
New
Old
Other
New
Old
Other
New
Old
Other
New
Old
Other
New
Old
Other
New
Old
Other
Llama3
MEMIT
21.5
36.5
42.0
16.8
50.3
32.9
29.7
43.8
26.5
15.6
46.9
37.5
21.0
37.5
41.5
48.0
27.0
25.0
RECT
19.3
39.1
41.5
15.2
51.6
33.2
26.5
46.4
27.1
11.7
49.7
38.5
16.9
42.3
40.8
44.0
30.1
25.9
PRUNE
21.0
35.8
43.2
17.8
49.3
32.9
31.1
42.6
26.3
12.6
49.5
37.9
23.0
37.7
39.3
44.9
28.7
26.4
AlphaEdit
24.6
33.8
41.6
20.7
47.3
32.0
30.0
42.6
27.4
17.1
43.1
39.7
24.9
35.6
39.5
54.3
23.2
22.5
CoME
25.5
32.9
41.6
19.8
48.3
32.0
29.0
44.1
26.9
18.5
42.7
38.8
29.8
33.1
37.1
56.1
21.1
22.8
TABLE XIII : Full breakdown of noncanonical tokenization responses across models, techniques, and datasets. New (%) reports the fraction producing the post-edit answer o∗ ; Old (%) reports the fraction recovering the pre-update answer o (bypass rate); Other (%) reports the fraction producing neither.
Open-weight language models publish the strings their chat templates use to mark turns, roles and tool results, which the tokenizer maps back to the reserved identifiers the model obeys. Anyone who controls text in a prompt can therefore write a turn boundary indistinguishable from one the serving stack wrote. We audit 256 deployed chat tokenizers. All are forgeable, and the flag usually recommended as a fix leaves 56.6% forgeable because it misses the tool and reasoning markers agent systems rely on. We propose nameless tokenization, which leaves the control entries with a reserved identifier and no surface string, so the content encoder cannot emit one and message content reaches the model unaltered. Across five tokenizer families it reproduces the standard token stream exactly on attack-free data and lifts accuracy on a probe of delimiter-bearing text from 8.5% to 59.9%, where sanitizers lose it. Separating a delimiter's appearance from its identifier shows the identifier matters little against a bare task instruction, but carries most of a forged tool result and most of any forged turn once the system message tells the model to treat user content as data.
Kisu Yang, Yoonna Jang, Heuiseok Lim
1VAIV Company · 2Hanwha Aerospace · 3Korea University
Machine unlearning has emerged as a critical capability for addressing privacy, safety, and regulatory concerns in large language models (LLMs). Existing methods operate at the sequence level, applying uniform updates across all tokens despite only a subset encoding the knowledge targeted for removal. This introduces gradient noise, degrades utility, and leads to suboptimal forgetting. We propose TokenUnlearn, a token-level attribution framework that identifies and selectively targets critical tokens. Our approach combines knowledge-aware signals via masking, and entropy-aware signals to yield importance scores for precise token selection. We develop two complementary strategies: hard selection, applying unlearning only to high-importance tokens, and soft weighting, modulating gradient contributions based on importance scores. Both extend existing methods to token-level variants. Theoretical analysis shows token-level selection improves gradient signal-to-noise ratio. Experiments on TOFU and WMDP benchmarks across three model architectures demonstrate consistent improvements over sequence-level baselines in both forgetting effectiveness and utility preservation.
Jiawei Wu, Doudou Zhou
Department of Statistics and Data Science, National University of Singapore.
Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversarial elicitation perspective, revealing that edited knowledge is often not fully erased and continues to surface, with consistent failures observed across diverse model architectures. To explain this behavior, we conduct a mechanistic analysis of popular KE methods. We show that low-rank updates do not overwrite existing knowledge but instead redistribute it within the model's representation space. Furthermore, we find that these methods act as targeted suppression mechanisms that reduce the likelihood of expressing original facts, rather than removing them from the model. Analysis of the loss landscape reveals that edited knowledge lies in narrow, anisotropic regions that are highly sensitive to perturbations, making them highly vulnerable to indirect prompting and adversarial attacks. By exposing these profound architectural vulnerabilities, our work proves that KE algorithms are inherently bypassable and motivates a fundamental reevaluation of how we deploy post-hoc updates in several LLM applications.
Advik Raj Basani, Anshuman Chhabra
Birla Institute of Technology and Science, Goa · University of South Florida