SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
Organizations: East China Normal University · Shanghai Innovation Institute · Xiamen University · University College London · Huawei Noah’s Ark Lab, UK · Independent Researcher · MemoraX AI
Abstract
LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.
Figures & tables
| Method | UOR | TSR | SUCR | |||
|---|---|---|---|---|---|---|
| Subset | Full | Subset | Full | Subset | Full | |
| DS-V4.1-Flash | 76.17% | 74.71% | 22.46% | 22.66% | 18.95% | 19.24% |
| DS-V4.1-Flash+Guard | 27.73% | 28.13% | 25.39% | 26.76% | 23.63% | 24.61% |
| SafeHarness | 41.67% | 40.38% | 43.85% | 43.65% | 31.55% | 30.76% |
| SHE | 28.77% | 24.06% | 58.71% | 63.93% | 48.14% | 54.35% |
| SafeCoEvo (Static) | 27.15% | 28.42% | 45.12% | 43.75% | 41.21% | 39.55% |
| Method | UOR | TSR | SUCR |
|---|---|---|---|
| DS-V4.1-Flash | 74.67% | 18.67% | 16.67% |
| DS-V4.1-Flash+Guard | 24.33% | 23.00% | 18.67% |
| SafeHarness | 27.68% | 62.11% | 53.86% |
| SHE | 29.19% | 54.36% | 41.95% |
| SafeCoEvo (Static) | 19.00% | 25.67% | 23.67% |
| SafeCoEvo w/o GuardVPO | 7.67% | 80.00% | 77.33% |
| Method | UOR | TSR | SUCR |
|---|---|---|---|
| DS-V4.1-Flash + SingGuard | 37.25% | 27.89% | 25.90% |
| SafeCoEvo w/o GuardVPO+SingGuard | 11.30% | 51.60% | 47.90% |
| Method | held-out Test Set | ||
|---|---|---|---|
| UOR | TSR | SUCR | |
| DS-V4.1-Flash+SingGuard | 39.67% | 26.00% | 23.00% |
| SafeCoEvo w/o GuardVPO+SingGuard | 7.33% | 49.33% | 46.33% |
| SafeCoEvo+SingGuard | 7.00% | 50.33% | 47.67% |
| Checkpoint | R-Judge | ATBench | |||||
|---|---|---|---|---|---|---|---|
| UnsafeF1 | UnsafeRec. | SafeSpec. | Correct | UnsafeF1 | UnsafeRec. | SafeSpec. | |
| AgentDoG 1.5 | 85.77% | 84.40% | 86.38% | 802/1000 | 79.33% | 72.75% | 89.58% |
| SafeCoEvo | 86.23% | 85.46% | 85.99% | 810/1000 | 80.61% | 75.51% | 88.35% |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | UOR | TSR | SUCR | |||
|---|---|---|---|---|---|---|
| Subset | Full | Subset | Full | Subset | Full | |
| DS-V4.1-Flash | 42.05% | 41.36% | 54.36% | 54.22% | 48.21% | 48.03% |
| DS-V4.1-Flash+Guard | 25.64% | 24.07% | 66.67% | 67.55% | 62.05% | 62.13% |
| SafeHarness | 41.18% | 40.24% | 63.10% | 62.15% | 50.27% | 49.52% |
| SHE | 39.18% | 40.65% | 63.92% | 64.26% | 51.03% | 49.2% |
| SafeCoEvo (Static) | 32.31% | 33.38% | 70.77% | 67.44% | 61.54% | 57.57% |
| Method | UOR | TSR | SUCR | |||
|---|---|---|---|---|---|---|
| Subset | Full | Subset | Full | Subset | Full | |
| DS-V4.1-Flash | 94.64% | 95.18% | 2.84% | 2.08% | 0.95% | 0.48% |
| DS-V4.1-Flash+Guard | 29.02% | 30.85% | 0.00% | 0.17% | 0.00% | 0.17% |
| SafeHarness | 41.96% | 40.45% | 32.49% | 32.09% | 20.50% | 19.00% |
| SHE | 22.40% | 13.02% | 55.52% | 63.9% | 46.37% | 58.01% |
| SafeCoEvo (Static) | 23.97% | 25.16% | 29.34% | 28.37% | 28.71% | 27.89% |
| Method | held-out Test Set | ||
|---|---|---|---|
| UOR | TSR | SUCR | |
| DS-V4.1-Flash | 45.00% | 52.00% | 49.00% |
| DS-V4.1-Flash+Guard | 29.00% | 69.00% | 56.00% |
| SafeHarness | 43.31% | 66.74% | 50.42% |
| SHE | 29.59% | 65.31% | 56.12% |
| SafeCoEvo (Static) | 17.00% | 73.00% | 67.00% |
| Method | held-out Test Set | ||
|---|---|---|---|
| UOR | TSR | SUCR | |
| DS-V4.1-Flash | 89.50% | 2.00% | 2.00% |
| DS-V4.1-Flash+Guard | 22.00% | 0.00% | 0.00% |
| SafeHarness | 16.18% | 58.71% | 56.39% |
| SHE | 29.00% | 49.00% | 35.00% |
| SafeCoEvo (Static) | 20.00% | 2.00% | 2.00% |
| Method | AgentDojo | AgentHarm | ||
|---|---|---|---|---|
| UOR | TSR | SUCR | ACC | |
| DS-V4.1-Flash | 54.72% | 16.98% | 0.00% | 12.41% |
| DS-V4.1-Flash+Guard | 0.00% | 28.30% | 28.30% | 44.53% |
| SafeHarness | 49.06% | 28.30% | 20.75% | 21.17% |
| SHE | 0.00% | 28.30% | 28.30% | 45.26% |
| SafeCoEvo (Static) | 0.00% | 20.80% | 20.80% | 45.26% |
| Setting | Method | UOR | TSR | SUCR |
| Safety | DS-V4.1-Flash+SingGuard | 24.86% | 69.73% | 65.95% |
| SafeCoEvo w/o GuardVPO+SingGuard | 20.50% | 73.30% | 65.60% | |
| Security | DS-V4.1-Flash+SingGuard | 44.48% | 3.47% | 2.52% |
| SafeCoEvo w/o GuardVPO+SingGuard | 5.70% | 38.20% | 36.90% |