cs.AISep 27, 2026

Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring

Authors: Zhixiang Zhang, Zesen Liu, Wai Ip Lai, Hongxu chen, Dongdong She

Organizations: The Hong Kong University of Science and Technology

Abstract

Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

    Aug 7, 2026Xiao Zhang, Yusheng Wang, Yuhao Fei +5Agent HarnessAttack-Success Rate

  2. SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

    Aug 10, 2026Wanying Qu, Qinghua Mao, Yu Li +12Large Language Model SafetyLarge Language Model Agents

  3. Auditing Agent Harness Safety

    May 14, 2026Chengzhi Liu, Yichen Guo, Yepeng Liu +8Agent HarnessModel Auditing