cs.AISep 29, 2026

SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

Authors: Yu Cheng, Yongkang Hu, Shuaijie Ma, Zhihang Lin, Weicheng Meng, Jingyang Qiao, Jiuan Zhou, Yushuo Zhang, +7 more

Organizations: East China Normal University · Shanghai Innovation Institute · Xiamen University · University College London · Huawei Noah’s Ark Lab, UK · Independent Researcher · MemoraX AI

Abstract

LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

    Aug 10, 2026Wanying Qu, Qinghua Mao, Yu Li +12Large Language Model SafetyLarge Language Model Agents

  2. Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

    Aug 13, 2026Xutao Mao, Liangjie Zhao, Xiang Zheng +1Large Language Model AgentsMalicious Agents

  3. Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution

    May 13, 2026Xiaozhe Zhang, Chaozhuo Li, Hui Liu +4Large Language Model SafetyModel-Agnostic Defense