Organizations: East China Normal University · Shanghai Innovation Institute · Xiamen University · University College London · Huawei Noah’s Ark Lab, UK · Independent Researcher · MemoraX AI
LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.
Figures & tables
Figure 1: From Repeated Evolution on Fixed Task Sets to Continual Test-Time Safety Adaptation.
Figure 2: SafeCoEvo: A Dual-Timescale Harness–Guard Co-Evolution Framework for Test-Time Agent Safety. S-Harness rapidly transforms recent runtime feedback into reusable explicit safety knowledge, while the Guard periodically internalizes accumulated event-level safety experience into parametric risk-judgment capabilities through GuardVPO.
Figure 3: S-Harness Evolution: From Runtime Feedback to Reusable Safety Knowledge.
Method
UOR ↓
TSR ↑
SUCR ↑
Subset
Full
Subset
Full
Subset
Full
DS-V4.1-Flash
76.17%
74.71%
22.46%
22.66%
18.95%
19.24%
DS-V4.1-Flash+Guard
27.73%
28.13%
25.39%
26.76%
23.63%
24.61%
SafeHarness
41.67%
40.38%
43.85%
43.65%
31.55%
30.76%
SHE
28.77%
24.06%
58.71%
63.93%
48.14%
54.35%
SafeCoEvo (Static)
27.15%
28.42%
45.12%
43.75%
41.21%
39.55%
Table 1: Overall performance on the continual evolution stream. Subset and Full denote the first 512 episodes and the full 1,024 episodes, respectively. Lower UOR and higher TSR/SUCR are better. Best and second-best results are shown in bold and underline, respectively.
Method
UOR ↓
TSR ↑
SUCR ↑
DS-V4.1-Flash
74.67%
18.67%
16.67%
DS-V4.1-Flash+Guard
24.33%
23.00%
18.67%
SafeHarness
27.68%
62.11%
53.86%
SHE
29.19%
54.36%
41.95%
SafeCoEvo (Static)
19.00%
25.67%
23.67%
SafeCoEvo w/o GuardVPO
7.67%
80.00%
77.33%
Table 2: Overall performance on the held-out test set.
Figure 6
Method
UOR ↓
TSR ↑
SUCR ↑
DS-V4.1-Flash + SingGuard
37.25%
27.89%
25.90%
SafeCoEvo w/o GuardVPO+SingGuard
11.30%
51.60%
47.90%
Table 3: Overall performance on the continual evolution stream with SingGuard as the Guard backbone.
Method
held-out Test Set
UOR ↓
TSR ↑
SUCR ↑
DS-V4.1-Flash+SingGuard
39.67%
26.00%
23.00%
SafeCoEvo w/o GuardVPO+SingGuard
7.33%
49.33%
46.33%
SafeCoEvo+SingGuard
7.00%
50.33%
47.67%
Table 4: Overall performance on the test set with SingGuard as the Guard backbone.
Checkpoint
R-Judge
ATBench
UnsafeF1
UnsafeRec.
SafeSpec.
Correct
UnsafeF1
UnsafeRec.
SafeSpec.
AgentDoG 1.5
85.77%
84.40%
86.38%
802/1000
79.33%
72.75%
89.58%
SafeCoEvo
86.23%
85.46%
85.99%
810/1000
80.61%
75.51%
88.35%
Table 5: Performance comparison on R-Judge and ATBench. F1 and Rec. denote unsafe-class F1 and recall, while Spec. denotes safe-class specificity.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Method
UOR ↓
TSR ↑
SUCR ↑
Subset
Full
Subset
Full
Subset
Full
DS-V4.1-Flash
42.05%
41.36%
54.36%
54.22%
48.21%
48.03%
DS-V4.1-Flash+Guard
25.64%
24.07%
66.67%
67.55%
62.05%
62.13%
SafeHarness
41.18%
40.24%
63.10%
62.15%
50.27%
49.52%
SHE
39.18%
40.65%
63.92%
64.26%
51.03%
49.2%
SafeCoEvo (Static)
32.31%
33.38%
70.77%
67.44%
61.54%
57.57%
Appendix
Table 6: Performance under the Safety setting on the continual evolution stream. Subset and Full denote the first 512 episodes and the full 1,024 episodes, respectively. Lower UOR and higher TSR/SUCR are better. Best and second-best results are shown in bold and underline, respectively.
Method
UOR ↓
TSR ↑
SUCR ↑
Subset
Full
Subset
Full
Subset
Full
DS-V4.1-Flash
94.64%
95.18%
2.84%
2.08%
0.95%
0.48%
DS-V4.1-Flash+Guard
29.02%
30.85%
0.00%
0.17%
0.00%
0.17%
SafeHarness
41.96%
40.45%
32.49%
32.09%
20.50%
19.00%
SHE
22.40%
13.02%
55.52%
63.9%
46.37%
58.01%
SafeCoEvo (Static)
23.97%
25.16%
29.34%
28.37%
28.71%
27.89%
Appendix
Table 7: Performance under the Security setting on the continual evolution stream. Subset and Full denote the first 512 episodes and the full 1,024 episodes, respectively. Lower UOR and higher TSR/SUCR are better. Best and second-best results are shown in bold and underline, respectively.
Method
held-out Test Set
UOR ↓
TSR ↑
SUCR ↑
DS-V4.1-Flash
45.00%
52.00%
49.00%
DS-V4.1-Flash+Guard
29.00%
69.00%
56.00%
SafeHarness
43.31%
66.74%
50.42%
SHE
29.59%
65.31%
56.12%
SafeCoEvo (Static)
17.00%
73.00%
67.00%
Appendix
Table 8: Test Safety Performance comparison on the held-out test set. All values are reported as percentages. ↓ indicates lower is better, while ↑ indicates higher is better. The best results are highlighted in bold, while the second-best results are underlined.
Method
held-out Test Set
UOR ↓
TSR ↑
SUCR ↑
DS-V4.1-Flash
89.50%
2.00%
2.00%
DS-V4.1-Flash+Guard
22.00%
0.00%
0.00%
SafeHarness
16.18%
58.71%
56.39%
SHE
29.00%
49.00%
35.00%
SafeCoEvo (Static)
20.00%
2.00%
2.00%
Appendix
Table 9: Test Security Performance comparison on the held-out test set. All values are reported as percentages. ↓ indicates lower is better, while ↑ indicates higher is better. The best results are highlighted in bold, while the second-best results are underlined.
Method
AgentDojo
AgentHarm
UOR ↓
TSR ↑
SUCR ↑
ACC ↑
DS-V4.1-Flash
54.72%
16.98%
0.00%
12.41%
DS-V4.1-Flash+Guard
0.00%
28.30%
28.30%
44.53%
SafeHarness
49.06%
28.30%
20.75%
21.17%
SHE
0.00%
28.30%
28.30%
45.26%
SafeCoEvo (Static)
0.00%
20.80%
20.80%
45.26%
Appendix
Table 10: Cross-benchmark performance on AgentDojo and AgentHarm. All values are reported as percentages. ↓ indicates lower is better, while ↑ indicates higher is better. The best results are highlighted in bold, while the second-best results are underlined.
Setting
Method
UOR ↓
TSR ↑
SUCR ↑
Safety
DS-V4.1-Flash+SingGuard
24.86%
69.73%
65.95%
SafeCoEvo w/o GuardVPO+SingGuard
20.50%
73.30%
65.60%
Security
DS-V4.1-Flash+SingGuard
44.48%
3.47%
2.52%
SafeCoEvo w/o GuardVPO+SingGuard
5.70%
38.20%
36.90%
Appendix
Table 11: Performance comparison under the safety and security settings. All values are reported as percentages. ↓ indicates lower is better, while ↑ indicates higher is better. The best result under each setting is highlighted in bold, and our method is shaded in light green.
Figure 6: Performance under the Safety setting on the continual evolution stream with Qwen3.7-Flash as the Target Agent. Lower UOR is better, while higher TSR and SUCR are better.
Figure 7: Performance under the Security setting on the continual evolution stream with Qwen3.7-Flash as the Target Agent. Lower UOR is better, while higher TSR and SUCR are better.
Figure 8: Performance under the Safety setting on the held-out test set with Qwen3.7-Flash as the Target Agent. Lower UOR is better, while higher TSR and SUCR are better.
Figure 9: Performance under the Security setting on the held-out test set with Qwen3.7-Flash as the Target Agent. Lower UOR is better, while higher TSR and SUCR are better.
Figure 10: Training dynamics and validation performance of GuardVPO. (a) Total loss and policy loss during optimization. (b) Validation unsafe outcome rate (UOR; lower is better). (c) Validation task success rate (TSR) and safe and successful completion rate (SUCR; higher is better). Validation metrics are reported as percentages.
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult. We propose Safety Harness Evolution (SHE), a framework that learns evolving safe boundaries from rollout trajectories. SHE decomposes the harness into four artifacts with explicit safety responsibilities, including the System Prompt, Rule Bank, Safety Memory, and Tool Policy, defining clear functional boundaries for localized evolution. Based on this decomposition, SHE introduces an attribution-guided evolution loop that converts trajectory failures into structured diagnoses, learns artifact-specific boundary refinements, and selects evolved harnesses through safety-utility validation. Experiments on Agent-SafetyBench demonstrate that SHE effectively enhances safety through harness evolution, achieving a 3.1x ASR reduction compared with static SafeHarness, while also improving benign utility. The evolved harness further generalizes to unseen risks on the held-out AgentHarm benchmark and transfers across agent models without additional evolution.
Wanying Qu, Qinghua Mao, Yu Li +12
Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · Fudan University +1
Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable procedures. Because evolution optimizes task outcomes rather than procedure safety, compromised experience can cause skill misevolution. Existing benchmarks measure current behavior or static artifacts but cannot attribute risk across authoring, retrieval, and later execution. To expose this lifecycle, we introduce SkillMisevo-Gym, a lifecycle-aware harness that versions skill state across agent frameworks, and SkillMisevo-Bench, a frozen design from malicious exposure to carryover tasks, with concept-aligned benign tasks and nine lifecycle metrics. We also introduce SafeEvolve, a wrapper that repairs unsafe content and governs subsequent reuse. Across 25 agent-method configurations, each covering 525 tasks in 25 episodes, all 21 evolved configurations author unsafe artifacts, while only fifteen lead to fresh-session harm. In the exposure sweep, three malicious tasks raise carryover ASR from 16.0% to 35.3%. Across representative skill evolution methods, SafeEvolve reduces unsafe retrieval and fresh-session harm by 26.7 and 17.3 percentage points, respectively, while mean benign utility changes by only 0.4 points. Together, persistent-adaptation safety must govern what updates write and what future executors reuse. Code is available at https://github.com/henrymao2004/misevolve.
Xutao Mao, Liangjie Zhao, Xiang Zheng +1
1City University of Hong Kong · 2Adelaide University
Large language models remain vulnerable to adversarial prompts that elicit harmful outputs. Existing safety paradigms typically couple red-teaming and post-training in a closed, policy-centric loop, causing attack discovery to suffer from rapid saturation and limiting the exposure of novel failure modes, while leaving defenses inefficient, rigid, and difficult to transfer across victim models. To this end, we propose EvoSafety, an LLM safety framework built around persistent, inspectable, and reusable external structures. For red teaming, EvoSafety equips the attack policy with an adversarial skill library, enabling continued vulnerability probing through simple library expansion after saturation, while supporting the evolution of adversarial vectors. For defense learning, EvoSafety replaces model-specific safety fine-tuning with a lightweight auxiliary defense model augmented with memory retrieval. This enables efficient, transferable, and model-agnostic safety improvements, while allowing robustness to be enhanced solely through memory updates. With a single training procedure, the defense policy can operate in both Steer and Guard modes: the former activates the victim model's intrinsic defense mechanisms, while the latter directly filters harmful inputs. Extensive experiments demonstrate the superiority of EvoSafety: in Guard mode, it achieves a 99.61% defense success rate, outperforming Qwen3Guard-8B by 14.13% with only 37.5% of its parameters, while preserving reasoning performance on benign queries. Warning: This paper contains potentially harmful text.
Xiaozhe Zhang, Chaozhuo Li, Hui Liu +4
City University of Hong Kong · Beijing University of Posts and Telecommunications · Wuhan University +2