Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker's subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model's preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.
Figures & tables
Figure 1 : Same bytes, different authority. The attacker controls the characters of a forged role marker inside a tool result. The server-side tokenizer decides whether these characters reach the model as one reserved control token or as ordinary subwords, and both encodings decode to identical text. Token boundaries are those of the Qwen3 tokenizer.
Condition
Forged markers
Construction
Same bytes as the forged payload
Reserved
reserved ids
the payload as the attacker wrote it, under default tokenization
Split
ordinary subwords
each marker encoded with the ordinary vocabulary, as under the standard mitigation
Matched
reserved ids
as Reserved , with ordinary text at the start of the tool response split so that the token count equals Split ’s
Different text
Plaintext
plain words
role labels written as plain words, such as System:
Table 1 : Encoding conditions used throughout the main text. The first block decodes to the same bytes and differs only in token ids; Plaintext is the same injection without template markers.
Model
Attack
Reserved
Matched
Split
Plaintext
Δ
Qwen3-8B
direct harm
84.8
83.4
75.3
20.4
+8.1
data stealing
90.9
90.0
89.5
41.6
+0.5
Llama-3.1-8B
direct harm
98.2
97.8
39.7
51.6
+58.2
data stealing
99.5
99.2
49.1
53.4
+50.2
GLM-4.5
direct harm
71.5
71.1
11.1
0.0
+60.0
data stealing
92.3
92.7
27.2
0.8
+65.5
Table 2 : Attack success rate (%) on InjecAgent and the identity gap Δ , Matched minus Split (pp). Rates are means over repeated runs on 400 paired cases per configuration. Bold gaps are significant in every run. Intervals and tests are given in App. B.1 .
Model
Total
Surface term
Qwen3-8B
+43.0
+39.1
Qwen3-32B
+31.1
+9.8
Llama-3.1-8B
+51.5
−3.4
GLM-4.5
+45.1
+4.1
Seed-OSS-36B
+55.1
+0.3
Table 3 : What the text of the marker is worth, direct harm, 400 cases, one experiment per model (pp). Total: Matched against a lookalike of the same length with no reserved id, such as <|zz_end|> . Surface term: Split against the same marker with its first letter upper-cased, which keeps its length, shape and token count. The Seed-OSS-36B surface term is not significant.
Vector at the marker position
Model
Reserved
Matched
Split
Mean
Nearest
Other reserved
Qwen3-8B
84.1
82.7
78.2
76.9
78.2
86.9
Llama-3.1-8B
98.2
97.6
39.8
58.4
98.4
96.1
Table 4 : Attack success (%) when only the input vector at each reserved marker position is replaced, direct harm, 510 cases. The replacement is the mean of the marker’s subword vectors, the vector of the nearest ordinary token in embedding space, or the vector of another reserved control token.
Defended
Model
Undefended
Six rules
Search
Best spelling
Qwen3-8B
85.7
74.9
81.2
closing | removed
Llama-3.1-8B
98.0
52.8
92.2
embedding neighbour
GLM-4.5
72.0
26.3
70.4
embedding neighbour
Seed-OSS-36B
84.8
33.0
72.6
embedding neighbour
Table 5 : Best attacker success (%) against the tokenizer-side defence, which encodes every reserved marker as ordinary subwords, direct harm, held-out cases. Undefended: the attacker’s best option without the defence. Six rules: the best of six respellings fixed in advance. Search: the best of 115 to 133 candidate spellings per model, selected on separate calibration cases. An embedding neighbour replaces each marker with an ordinary token close to it in input-embedding space.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Truncated (%)
Δ (pp)
Model
Attack
Matched
Split
All pairs
Untruncated
Pairs kept (%)
Qwen3-8B
DH
3.7
2.9
+8.1
+8.4
94
DS
5.7
5.5
+0.5
+0.7
89
Llama-3.1-8B
DH
0.0
0.0
+58.2
+58.2
100
DS
1.0
3.0
+50.2
+51.6
96
GLM-4.5
DH
0.0
0.0
+60.0
+60.0
100
Appendix
Table 6 : Truncation and the identity gap. The last two columns restrict each configuration to pairs in which neither Matched nor Split was truncated.
Model
Reserved
Plaintext
Split
Qwen3-8B
0–1
0–1
16–18
Llama-3.1-8B
0
0
31–33
Appendix
Table 7 : Extra tokens of the span-wise prompt over tokenizing the whole rendered prompt, over 120 cases per model.
Condition
Construction
Used in
Same bytes, reserved ids kept
Reserved
the payload as the attacker wrote it, under default tokenization
throughout
Matched
as Reserved , with ordinary text at the start of the tool response split so that the token count equals Split ’s
throughout
Matched-before
as Matched , but splitting the ordinary text that ends at the marker
§ 4.3 , § 4.5
Matched-after
as Matched , but splitting the ordinary text that starts after the marker
§ 4.3 , § 4.5
Char-matched
as Matched , at Char-split ’s token count
§ 4.3
Appendix
Table 8 : All encoding conditions, grouped by whether they keep the bytes of the forged payload and its reserved ids.
Check
Rules out
Decoded prompt equals the intended string
any byte changed by the encoder
Reserved and Split decode to identical bytes
a content difference between conditions
Split ’s untrusted span contains no reserved id
a split condition that is still partly reserved
Conditions agree outside the untrusted span
merges across the span boundary
Reserved ’s span contains reserved ids
markers that the tokenizer does not reserve
Matched ’s span contains reserved ids
Matched being a second copy of Split
Appendix
Table 9 : Construction checks applied to every case.
Model
Attack
Δ (pp)
95% interval
Largest p
Runs
Extra-token cost
Qwen3-8B
DH
+8.1
[+2.2,+14.2]
5.6×10−3
5
+1.4
DS
+0.5
[−6.2,+5.8]
0.80
5
+0.9
Llama-3.1-8B
DH
+58.2
[+53.0,+63.3]
<10−10
3
+0.3
DS
+50.2
[+45.0,+55.5]
<10−10
3
+0.2
GLM-4.5
DH
+60.0
[+54.8,+65.2]
<10−10
3
+0.4
DS
+65.5
[+59.2,+71.7]
<10−10
3
−0.3
Appendix
Table 10 : Identity gap with its statistics. The interval is the envelope of the per-run paired bootstrap 95% intervals, p is the largest exact McNemar p -value over runs, and the last column is the cost of the extra tokens alone, the success rate of Reserved minus that of Matched .
Model
Attack
Primary
Estimates
Range
Qwen3-8B
DH
+8.1
16
+4.5 to +12.0
DS
+0.5
7
+0.3 to +1.1
Llama-3.1-8B
DH
+58.2
12
+52.0 to +58.6
DS
+50.2
5
+50.0 to +55.3
GLM-4.5
DH
+60.0
7
+57.8 to +60.4
DS
+65.5
5
+65.5 to +66.6
Appendix
Table 11 : Range of Δ across the InjecAgent experiments that measure it with the original forged block and the reasoning block on (pp). The first column is the estimate of Table 2 , or of Table 17 for Qwen3-32B.
Model
Attack
Cost
Interval
Margin
Equivalent to zero
Qwen3-8B
DH
+1.4
[−3.0,+6.0]
4.8
no
DS
+0.9
[−3.5,+6.3]
0.7
no
Llama-3.1-8B
DH
+0.3
[−1.0,+1.8]
29.3
yes
DS
+0.2
[−0.8,+1.3]
25.3
yes
GLM-4.5
DH
+0.4
[−2.0,+3.0]
30.2
yes
DS
−0.3
[−3.0,+2.3]
32.5
yes
Appendix
Table 12 : Equivalence test for the cost of the extra tokens (pp). The interval resamples runs and cases.
Attack
Cases
Reserved
Matched
Split
Δ
Extra-token cost
DH
510
85.2
82.9
75.8
+7.1
+2.4
DS
543
90.7
90.7
89.6
+1.0
0.0
Appendix
Table 13 : Qwen3-8B on every case of the benchmark: success rates (%) and gaps (pp).
Model
Attack
Draw 1
Draw 2
Draw 3
Qwen3-8B
DH
+8.1
+4.9
+6.9
DS
+0.5
+0.3
+0.9
Llama-3.1-8B
DH
+58.2
+58.3
+57.8
DS
+50.2
+54.3
+55.3
GLM-4.5
DH
+60.0
+58.6
+60.0
DS
+65.5
+65.8
+66.4
Appendix
Table 14 : Identity gap Δ (pp) on three case draws of 400 cases.
Model
Attack
Primary
Required arguments
Qwen3-8B
DH
+5.5
+5.8
DS
+0.3
+0.3
Llama-3.1-8B
DH
+58.2
+51.8
DS
+53.3
+42.6
GLM-4.5
DH
+60.4
+60.4
DS
+65.5
+65.5
Appendix
Table 15 : Identity gap (pp) in the experiment of Table 22 under the primary criterion and under the stricter criterion that also requires every required argument.
Model
transformers
vLLM
Llama-3.1-8B
+55.0
+55.2
Seed-OSS-36B
+55.0
+55.7
Qwen3-8B
+8.0
+12.0
Qwen3-32B
+15.0
+18.3
Appendix
Table 16 : Identity gap (pp) under two inference engines on identical cases, direct harm.
Model
Attack
Reserved
Matched
Split
Plaintext
Perturbed
Δ
Qwen3-32B
DH
88.0
90.0
72.0
37.6
42.5
+18.0
DS
95.9
96.1
92.8
75.0
55.1
+3.3
Seed-OSS-36B
DH, 1536
78.8
78.2
23.8
27.3
18.8
+54.4
DH, 4096
85.0
84.8
26.2
31.1
21.4
+58.6
Appendix
Table 17 : Success rates (%) and identity gaps (pp) for Qwen3-32B, and for Seed-OSS-36B at token budgets of 1536 and 4096. Perturbed is the template with 10% of its characters altered (Table 8 ).
Model
Forged block
Reserved
Matched
Split
Δ
Qwen3-8B
original
84.8
82.3
76.8
+5.5
acknowledge, then user
94.1
94.4
94.7
−0.3
system, then assistant
91.4
89.3
70.5
+18.8
Llama-3.1-8B
original
98.2
97.8
39.6
+58.2
acknowledge, then user
100.0
100.0
45.4
+54.6
system, then assistant
99.0
99.0
37.5
+61.5
Appendix
Table 18 : Success rates (%) and identity gaps (pp) for three forged blocks, direct harm, every case of the benchmark, three runs. The rows for the original block come from the experiment of Table 22 .
Model
Attack
User turn
System turn
Qwen3-8B
DH
+8.3
+9.1
DS
−1.6
+1.1
Llama-3.1-8B
DH
+55.2
+57.7
DS
+48.3
+50.0
GLM-4.5
DH
+34.1
+59.0
Appendix
Table 19 : Identity gap (pp) with a forged user turn and a forged system turn, measured together.
Subword split,
Subword split,
Per-character split,
Model
Attack
count-matched ( Δ )
position-matched
count-matched
Qwen3-8B
DH
+8.1
+11.4
+54.2
DS
+0.5
+3.8
+28.3
Llama-3.1-8B
DH
+58.2
+43.3
+37.3
DS
+50.2
+42.1
+33.7
GLM-4.5
DH
+60.0
+59.4
+66.2
Appendix
Table 20 : Identity gap (pp) for three combinations of split rule and matched control. The position-matched column comes from a separate experiment with Matched-after alone; Table 22 gives the experiment with all three placements. On Llama-3.1 the per-character control is not null (see text), so that column is a lower bound there.
Comparison
Qwen3-8B
Llama-3.1-8B
Char-matched against Char-split (per-character gap)
+54.3
+37.8
Matched against Split (identity gap Δ )
+7.5
+58.6
Reserved against Char-matched (cost of extra tokens)
+2.2
+9.4
Char-split against Split
−46.6
+11.4
Appendix
Table 21 : Subword and per-character splitting in one experiment, direct harm (pp).
Gap to Split
Cost of extra tokens
Model
Attack
Matched
Before
After
Before
After
Qwen3-8B
DH
+5.5
+6.3
+9.9
+1.7
−1.9
DS
+0.3
+0.1
+4.1
+0.6
−3.4
Llama-3.1-8B
DH
+58.2
+59.5
+44.1
−0.8
+14.5
DS
+53.3
+53.9
+46.0
−0.2
+7.7
GLM-4.5
DH
+60.4
+60.5
+58.1
−0.7
+1.7
Appendix
Table 22 : Position-matched controls on every case of the benchmark (pp). The first three columns give the gap between each matched condition and Split ; the last two give the cost of the extra tokens, Reserved minus the matched condition. Before and after denote Matched-before and Matched-after .
Model
Total
Identity
Proximity
Identity share
Qwen3-8B
+55.1
+42.8
+12.3
78%
Qwen3-32B
+35.1
+18.3
+16.8
52%
Llama-3.1-8B
+44.7
+46.8
−2.1
105%
Appendix
Table 23 : Identity and embedding proximity at equal bytes, position and token count, direct harm, 400 cases, three runs (pp). The total is the gap between Emb-matched and Emb-far ; the identity share is the fraction of the total carried by the identity term.
Model
Cases
Reserved
Char-split
Drop (pp)
P>0.99
Qwen3-8B
209 of 400
99.998
99.860
0.14
100%
Llama-3.1-8B
184 of 400
95.52
55.72
39.80
48.9%
Appendix
Table 24 : The prior readout on our data: probability of the target call (%) on the doubly conditioned subsample, and the share of conditioned cases on which this probability exceeds 0.99 under Reserved . The subsample keeps the direct-harm cases on which Reserved succeeds and Plaintext fails in every run of Table 2 .
All cases
Condition 1
Conditions 1 and 2
Model
Attack
Δ
Δ
Cases
Δ
Cases
Qwen3-8B
DH
+8.1
+9.7
339
+12.8
263
DS
+0.5
+0.6
363
+1.7
208
Llama-3.1-8B
DH
+58.2
+58.8
392
+90.6
187
DS
+50.2
+49.9
398
+75.4
184
GLM-4.5
DH
+60.0
+81.2
286
+81.2
286
Appendix
Table 25 : Identity gap (pp) under the sample conditioning of Deng et al. (2026) . Condition 1 keeps cases where Reserved succeeds; condition 2 further requires that Plaintext fails. Each run is conditioned on its own outcomes and the gap is averaged over runs; Cases is the average count per run. Table 24 instead keeps only the cases that meet both conditions in every run.
Model
Quantity
Instruct
Base
Shift
Wilcoxon p
Qwen3-1.7B
identity gap
+7.82
+0.15
+7.67
3.2×10−25
extra-token cost
+0.45
+0.08
+0.37
0.11
Qwen3-8B
identity gap
+0.68
−0.46
+1.14
4.6×10−7
extra-token cost
+0.21
+0.11
+0.10
0.20
Seed-OSS-36B
identity gap
+0.33
−2.60
+2.92
3.4×10−29
extra-token cost
+0.11
+0.01
+0.10
0.15
Appendix
Table 26 : Identity gap and extra-token cost in logits for base and instruction-tuned checkpoints, 200 cases.
Attack
Condition
Reasoning on
Reasoning off
DH
Reserved
84.3
86.8
Matched
83.5
85.5
Split
74.7
35.7
identity gap Δ
+8.8
+49.8
DS
Reserved
90.8
96.2
Matched
89.9
94.9
Appendix
Table 27 : Reasoning suppression on Qwen3-8B with both settings run together: success rates (%) and identity gaps (pp).
Model
Suite
Pairs
Matched
Matched-before
Matched-after
Qwen3-8B
all four
409
+10.3
+9.5
+8.2
banking
89
+2.2
+2.7
+9.0
slack
50
+14.0
+7.4
+5.3
travel
85
+27.4
+25.6
+14.5
workspace
185
+5.2
+6.6
+5.8
Qwen3-32B
all four
409
+6.8
+6.2
+6.4
Appendix
Table 28 : AgentDojo, four-suite held-out split: gap between each matched condition and Split (pp).
Model
Suite
Pairs
Reserved
Matched
Split
Δ
Qwen3-8B
banking
89
21.0
23.6
18.4
+5.2
slack
50
41.3
42.0
33.3
+8.7
travel ⋆
85
26.3
30.2
4.7
+25.5
Llama-3.1-8B
banking ⋆
89
22.5
19.1
6.7
+12.4
Qwen3-32B
banking
89
17.6
21.7
13.5
+8.2
slack
50
24.7
26.7
6.7
+20.0
Appendix
Table 29 : AgentDojo, three-suite held-out split: success rates (%) and identity gaps (pp). Stars mark the two primary tests fixed in advance.
Checkpoints
Count
Left intact
Tool-protocol
DeepSeek-V3.1
1
18
13
GLM-4.5, GLM-4.6
2
14
10
Kimi-K2-Thinking
1
7
7
Qwen3 (8B, 14B, 30B-A3B, 235B-A22B, …)
7
12
6
Qwen3.5-4B, Qwen3.5-9B
2
12
6
Seed-OSS-36B-Instruct
1
6
6
Appendix
Table 30 : Added tokens that split_special_tokens leaves intact, and the subset in the tool-protocol class, which includes reasoning delimiters, for the families studied here and their close relatives.
Model
Tool channel
System channel
Reserved : tool minus system
Qwen3-8B
+11.7
+8.4
−25.2
Qwen3-32B
+19.9
+17.7
−18.0
GLM-4.5
+15.7
+60.2
−50.7
Seed-OSS-36B
+9.4
+58.8
−45.2
Appendix
Table 31 : Identity gap (pp) for a forged block of tool-protocol tokens, which the standard mitigation leaves intact, and for the original forged system turn, whose tokens it covers, direct harm, one experiment per model. The last column is the change in Reserved success when the tool block replaces the system block.
Defended best
Removed
Model
Attack
Best spelling
Undefended
Six rules
Search
Six rules
Search
Qwen3-8B
DH
closing | removed
85.7
74.9
81.2
+10.8
+4.5
Llama-3.1-8B
DH
neighbour, rank 14
98.0
52.8
92.2
+45.2
+5.8
GLM-4.5
DH
neighbour, rank 9
72.0
26.3
70.4
+45.7
+1.6
Seed-OSS-36B
DH
neighbour, rank 8
84.8
33.0
72.6
+51.8
+12.2
Qwen3-8B
DS
closing | removed
93.1
89.2
93.1
+1.2
0.0
Appendix
Table 32 : Best attacker success (%) with and without the defence, and the success the defence removes (pp). A neighbour of rank k replaces each marker with the k -th nearest ordinary token to the reserved one in input-embedding space. Undefended is the attacker’s best option without the defence. For the six-rule columns, the undefended baseline is the best of Reserved and the six rules, which differs from the Undefended column only for Qwen3-8B DS ( 90.4% ).