Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks
Authors: Emanuele La Malfa, Saar Cohen, Gabriele La Malfa, Mickel Liu, Christian Schroeder de Witt, Natasha Jaques, Michael J. Wooldridge
Organizations: Department of Computer Science, University of Oxford · Institute for Decentralized AI (IDAI) · Department of Informatics, King’s College London · Paul G. Allen School of Computer Science & Engineering, University of Washington · UCL Computer Science & AI Centre, University College London
Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show that no fixed bounded window of recent prompts is sufficient in general: safety-relevant information may occur arbitrarily far back in the interaction. We formalize a watchman, an online state mechanism that carries this information forward, and show that under explicit assumptions it enables zero-failure defense with positive benign helpfulness. Under stronger conditions, it is also optimal among zero-failure defenders. An exact watchman may nevertheless require exponentially many states, while exact maliciousness detection can require exponentially many queries in an unstructured black-box model. These state and query lower bounds do not by themselves imply hard learning: the construction underlying the state lower bound is efficiently learnable from labeled examples, whereas certifying worst-case safety can require substantially more information under restricted access. We also show that self-play equilibrium alone does not certify usefulness, motivating a constrained formulation that maximizes worst-case benign helpfulness among zero-failure defenders. Empirically, training role-specific attacker and defender LoRA adapters over frozen LLMs via multi-turn self-play strengthens both roles: attackers become more effective at eliciting harmful responses, while defenders become more robust to attack, with improvements also observed on unseen attack objectives.
Figures & tables
Figure 1: (A) A mosaic attacker decomposes a harmful objective into individually benign subprompts. Each response can be appropriate in isolation, while the collection of responses becomes harmful through composition. (B) Unlike escalation attacks such as Crescendo ( Russinovich et al., 2025 ) , mosaic fragments need not become progressively more harmful; their risk arises from composition. (C) We train role-specific LoRA adapters over a shared frozen backbone; a judge scores each conversation to update both roles through self-play ( Liu et al., 2026 ; La Malfa et al., 2026 ) .
Model
Benchmark
Base Model
Attacker vs Base ( ↑ )
Base vs Defender ( ↓ )
Benign refusal (pp)
ASR
ASR
Δ (pp)
ASR
Δ (pp)
Base
Def.
Δ
Qwen3.6-35B-A3B
WildMosaic (Held-out)
0.170
0.290 ( 0.020 )
+12.0
0.077 ( 0.021 )
+9.3
0.0
0.3 ( 0.6 )
+0.3
AdvBench
0.224 ( 0.016 )
0.264 ( 0.010 )
+4.0†
0.071 ( 0.016 )
+15.3
1.9 ( 0.2 )
7.6 ( 0.4 )
+5.7
WildJailbreak
0.154 ( 0.013 )
0.178 ( 0.016 )
+2.4
0.073 ( 0.010 )
+8.1
1.4 ( 0.9 )
6.7 ( 0.4 )
+5.3
Gemma4-26B-A4B
WildMosaic (Held-out)
0.630
0.743 ( 0.012 )
+11.3
0.485 ( 0.058 )
+14.5
0.0
5.8 ( 5.1 )
+5.8
AdvBench
0.699 ( 0.019 )
0.803 ( 0.005 )
+10.4
0.525 ( 0.035 )
+17.4
0.0 ( 0.0 )
5.5 ( 5.3 )
+5.5
Table 1: Attack success rate (ASR) on the WildMosaic held-out split, AdvBench, and WildJailbreak. Base is the frozen model’s ASR; Attacker is the LoRA attacker against the base defender ( ↑ better), and Defender is the base attacker against the LoRA defender ( ↓ better). Δ is the absolute ASR change in percentage points (pp), i.e. ASRatt−ASRbase for the attacker and ASRbase−ASRdef for the defender. Values are mean with standard deviation in parentheses across three evaluated checkpoints. For AdvBench and WildJailbreak, significance is assessed separately at each checkpoint using an exact two-sided McNemar test on paired malicious prompts ( n=520 for Qwen and Gemma and n=300 for DeepSeek; α=0.05 ); checkpoint-level tests are not combined into a test of the reported mean. Bold Δ indicates significance at all three evaluated checkpoints, while † indicates significance at one or two checkpoints. No significance marker is reported for WildMosaic, for which checkpoint-level McNemar tests are not reported here. Benign refusal is the share of benign prompts refused by the base model and the trained defender and is reported separately to characterize the safety-helpfulness trade-off.
Figure 2: Jailbreak rate by harm topic: base model vs. self-play defender. Each bar shows how often a harmful prompt jailbroke the model. Grey for the frozen base model, blue for the self-play defender, grouped by the topic of the harmful request, on AdvBench (left) and WildJailbreak (right). The labels were assigned to one of eight fixed harm categories using an LLM classifier (Qwen3.6-35B-A3B, constrained to the fixed label set). Across nearly every topic on both benchmarks and both models, the defender lowers the jailbreak rate.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen
Gemma
DeepSeek
Base model
Qwen3.6-35B-A3B
Gemma-4-26B-A4B-it
DeepSeek-V2-Lite-Chat
Phases
2: steps 0 → 40 from grpo-atk1.5e-4 , then 40 → 65 as fullgrpo-beta0-resume40
2: steps 0 → 40 from a2d5-g20-kl0.03-dlr1.5e-4-to40 , then 40 → 80
1: steps 0 → 90 in a single run
Training steps
65
80
90
G (rollouts per seed)
10 in both phases
20 in both phases
20
β (KL coefficient, k3)
P1: 0.03 → P2: 0 (no reference model)
P1: 0.03 → P2: 0 (no reference model)
0.03
top- k (A / D)
P1: 2 / 2 → P2: 5 / 5
P1: 2 / 5 → P2: 5 / 5
3 / 3
Appendix
Table 2: Training configuration for the headline adversarial self-play runs used in our experiments. Additional runs and configurations contributing evaluated checkpoints are described in the text.
Model
Benchmark
Base
Attacker ( ↑ )
Defender ( ↓ )
Attacker vs. Defender
Benign refusal
ASR
ASR
Δ
ASR
Δ
ΔA
ΔD
base → def. (%)
Qwen3.6-35B-A3B
AdvBench
0.208
0.273
+31.5%
0.085
+59.3%
+40.9%
+56.3%
1.7 → 7.9
WildJailbreak
0.163
0.196
+20.0%
0.079
+51.8%
+24.4%
+50.0%
2.5 → 6.9
Pooled
0.186
0.235
+26.4%
0.082
+56.0%
+32.9%
+53.7%
2.1 → 7.4
Appendix
Table 3: Four-condition PAIR-adapted attack success rate for Qwen3.6-35B-A3B. Base : base attacker vs. base defender. Attacker : abs -A vs. base defender ( ↑ better). Defender : base attacker vs. abs -D ( ↓ better). Attacker vs. Defender : abs -A vs. abs -D; ΔA is relative to Defender and ΔD to Attacker . Each Δ is the relative improvement in the indicated direction. Pooled combines the two disjoint benchmark goal sets. Bold Δ indicates significance (exact McNemar test, α=0.05 ). One deterministic run is used per condition. Benign refusal reports the refusal rate on benign prompts, base → abs -D.
parameterization
Prompts
n
cos
95% CI
z
p
frac. neg.
Base, shared backbone
all
500
−0.00831
[−0.01000,−0.00663]
−2.18
3.2×10−20 ***
0.66
malicious
250
−0.00998
[−0.01267,−0.00733]
−2.53
4.6×10−12 ***
0.66
benign
250
−0.00663
[−0.00867,−0.00458]
−1.84
8.7×10−10 ***
0.67
abs , projected onto backbone
all
499
−0.00622
[−0.00776,−0.00470]
−1.49
1.0×10−14 ***
0.65
malicious
249
−0.00505
[−0.00734,−0.00276]
−1.04
2.6×10−5 ***
0.60
benign
250
−0.00738
[−0.00939,−0.00542]
−1.92
5.3×10−12 ***
0.70
Appendix
Table 4: Per-episode cosine similarity between the attacker’s and the defender’s policy gradient, Qwen3.6-35B-A3B at checkpoint fullgrpo-beta0-resume40 step 65. 250 AdvBench goals and 250 benign prompts, identical seeds in every condition. Negative cosine means the two updates conflict. z=cos⋅deff expresses the cosine in units of its chance scale; p is a two-sided one-sample t -test of the per-episode cosines against zero (*** p<10−3 ); CI is a bootstrap 95% interval on the mean. frac. neg. is the fraction of episodes with a negative cosine ( 0.5 under the null hypothesis).
Model
Base Model ASR
Defender ASR
Δ
p
Qwen3.6-35B-A3B
0.28
0.27
−0.01
0.82
Gemma4-26B-A4B
0.73
0.61
−0.12
0.003
DeepSeek-V2-Lite
0.065
0.058
−0.008
0.85
Appendix
Table 5: Transfer of self-play red-teaming to Crescendo, an unseen multi-turn jailbreak. Attack success rate (ASR) of the base model vs. the self-play defender, each model attacked by its own base on identical prompts; lower ASR is better for the defender. Δ is the absolute change and p is derived from an exact McNemar test over n=260 paired malicious conversations. Bold marks a significant reduction ( α=0.05 ).
Feature
Qwen
Gemma
DeepSeek
SD base
SD train
Disp. ↓
Seed-word recall, payload turn (%)
25.6 → 12.3
14.9 → 15.7
12.1 → 13.8
0 5.8
0 1.4
4.2 ×
Attacker words per episode
176 → 172
132 → 132
0 98 → 128
32.0
20.1
1.6 ×
Chaining cue, turns 2+ (%)
49.7 → 41.9
33.8 → 35.6
18.8 → 23.4
12.6
0 7.7
1.6 ×
Specificity push, turns 2+ (%)
83.3 → 66.1
69.8 → 77.1
36.2 → 47.3
19.8
12.3
1.6 ×
Cover story, turn 1 (%)
38.7 → 33.7
21.5 → 22.1
12.5 → 16.7
10.9
0 7.1
1.5 ×
Turn-1 length (words)
43.7 → 43.9
24.0 → 25.7
22.0 → 30.9
0 9.8
0 7.7
1.3 ×
Appendix
Table 6: Attacker linguistic features on WildJailbreak. SD is the cross-family standard deviation; the last column is the reduction in dispersion.
Family
Ckpt
Refusal (base)
Refusal (trained)
Δ tokens
D/A base
D/A trained
Qwen
55
27 / 43 / 54
42 / 62 / 74
−20.0%
12.3
0 10.0
Gemma
65
0 3 / 0 4 / 0 7
0 2 / 0 8 / 22
0 +0.2%
17.8
18.0
DeepSeek
80
13 / 0 5 / 0 2
18 / 11 / 0 9
0 −6.7%
0 7.7
0 7.1
Appendix
Table 7: Defender behavior facing the base attacker (WildJailbreak). Refusal rate (%) at turns 1/2/3, change in cumulative defender tokens on malicious episodes, and leverage D/A (defender tokens per attacker token).
Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities. This sequential setup leads to attackers overfitting obsolete exploits while defenders perpetually lag behind emerging threats. To address this, we introduce Self-RedTeam, the first fully online self-play multi-agent reinforcement learning (MARL) algorithm that continuously co-evolves attacker and defender for robust safety alignment. A single policy self-plays as both attacker and defender, generating adversarial prompts and defending against them, with a reward model adjudicating outcomes. Each role uses hidden chain-of-thought for strategic planning. Grounded in two-player zero-sum game theory, we establish a theoretical safety guarantee: if the game converges to Nash Equilibrium, the defender produces safe responses against any adversarial input. Empirically, Self-RedTeam generalizes across five models from the Llama and Qwen families, uncovering more diverse attacks (+17.80% SBERT) and improving safety of RLHF-trained models by up to 95% across 14 benchmarks. Our work motivates a shift from reactive patching to proactive co-evolution, enabling LLM safety self-improvement via online self-play MARL. Link to code: https://github.com/mickelliu/selfplay-redteaming
Mickel Liu, Liwei Jiang, Yancheng Liang +4
Paul G. Allen School of Computer Science & Engineering, University of Washington, Seattle, WA, USA · Department of Computer Science, Stanford University, Stanford, CA, USA
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100% at only a 5% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.
Yibo Zhang, Tianrong Guan, Liang Lin +3
Queen Mary University of London, UK · Squirrel AI Learning, USA
Large language model (LLM) safety classifiers such as Llama Guard are effective at detecting overtly harmful prompts but remain vulnerable to adversarial jailbreak attacks that disguise malicious intent through role-play scenarios, fictional framing, and indirect requests. We present Reflect-Guard, a method that augments LLM-based safety classifiers with chain-of-thought self-reflection capabilities through parameter-efficient fine-tuning. Our approach distills analytical reasoning from GPT-4o-mini into structured reflection annotations, then trains Llama-Guard-3-8B via QLoRA to generate logical self-reflections before issuing safety verdicts. Using only 1000 training examples and updating just 0.5% of model parameters (~42M), Reflect-Guard achieves substantial improvements on two challenging benchmarks. On WildGuardTest, F1 score improves from 0.770 to 0.842 (+7.2 pp), with recall on adversarial prompts increasing from 0.513 to 0.921 (+40.8 pp). On JailbreakBench, the attack success rate drops from 10.3% to 1.8%, representing an 82.5% relative reduction. These gains are especially pronounced on adversarial inputs, where the explicit reasoning step enables the model to see through obfuscation techniques that defeat standard pattern-matching approaches. Our results demonstrate that teaching safety classifiers to reason about adversarial intent, rather than simply classify surface patterns, is a promising direction for robust LLM safety.
Lixing Lin, Juli You, Yue Li +4
yaleYale University, New Haven, CT, USA · indIndependent Researcher · columbiaColumbia University, New York, NY, USA +2