Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.
Figures & tables
Figure 1: A motivating example of compositional safety failure in harness evolution. The arrows denote two independently accepted updates: the prompt update P0→P1 adds Bob to authorize_group , while the memory update M0→M1 adds BoA to bank_account . Under the rule that authorize_group can access bank_account , either update alone remains safe: Bob is not authorized under P0 , and BoA is unavailable under M0 . However, when the updated states P1 and M1 coexist, Bob can access BoA, producing a compositional safety failure.
Method
Phase
Verification Scope
Interaction-Level Analysis
TTHE ( Nie et al., 2026 )
Test-time
Candidate harness
✗
AutoSaddler ( Park et al., 2026 )
Offline
Component patch
✗
HarnessLens ( Xu et al., 2026a )
Offline
Harness modification
✗
SHE ( Qu et al., 2026 )
Offline
Attributed update
✗
SafeEvolve ( Mao et al., 2026 )
Offline
Candidate harness
✗
Ours
Runtime
Component-state interaction
✓
Table 1: Position of our work with existing literature. Existing works mainly focus on individual components, while our work analyzes the cross-component interactions.
Benchmark
Panels
Eligible
CF
CFR
AgentDojo ( Debenedetti et al., 2024 )
600
157
22
14.01%
Agent-SafetyBench ( Zhang et al., 2024b )
182
11
2
18.18%
Agent Security Bench ( Zhang et al., 2025 )
1362
1243
19
1.53%
Table 2: Results of pairwise compositional safety failure identification.
Figure 2: Overview of our method. (a) A candidate update ut introduces a component-state node ( t+ in this example), and the local interaction neighborhood Γh(ut) is extracted from the existing hypergraph Gt . (b) Local completion proposes interactions involving the new state. An LLM filters these candidates using component states and provenance, yielding the provisional hypergraph Gt+1 . Solid and dashed contours denote observed and predicted hyperedges, respectively; node shapes encode component types. (c) Within one action context, participating nodes become active (filled) in sequence. Each activation visits only incident hyperedges; a fully activated interaction is checked before the action takes effect. Execution provenance promotes a predicted hyperedge to observed status once confirmed at runtime, irrespective of its safety verdict.
Method
AgentDojo
Agent-SafetyBench
Agent Security Bench
No Monitoring
22/157 (14.01%)
2/11 (18.18%)
19/1243 (1.53%)
Naive
1/157 (0.64%)
0/11 (0.00%)
0/1243 (0.00%)
SHE ( Qu et al., 2026 )
6/157 (3.82%)
1/11 (9.09%)
5/1243 (0.40%)
HarnessLens ( Xu et al., 2026a )
8/157 (5.10%)
1/11 (9.09%)
5/1243 (0.40%)
Ours
2/157 (1.27%)
0/11 (0.00%)
1/1243 (0.08%)
Table 3: Safety effectiveness of different mechanisms against compositional failures. Each entry reports the residual compositional failure rate (CFR) as f/n ( % ), where f is the number of compositional failures among the n eligible panels. The eligible panel set is fixed before applying each safety mechanism.
Method
Utility ↑
Verification ↓
Runtime Check Rate ↓
Time Cost (s/task) ↓
AgentDojo
Naive
58.6%
–
100.0%
103.9
SHE ( Qu et al., 2026 )
57.3%
1200
–
77.1
HarnessLens ( Xu et al., 2026a )
52.9%
114
–
33.8
Ours
59.2%
82
33.6%
32.5
Agent-SafetyBench
Table 4: Utility and verification cost of different safety mechanisms. Utility is the fraction of eligible panels satisfying the benchmark-specific task objective. Verification counts additional validation executions per harness update. Runtime Check Rate is the fraction of candidate actions invoking the runtime safety checker. Time Cost is the additional wall-clock time per panel over execution without a safety mechanism.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Triples
Eligible
3-CF
3-CFR
AgentDojo ( Debenedetti et al., 2024 )
540
138
15
10.87%
Agent Security Bench ( Zhang et al., 2025 )
681
578
3
0.52%
Appendix
Table 5: Identification of 3-way compositional failures. A triple is eligible only when the baseline, all individual updates, and all pairwise compositions are safe and utility-preserving. 3-CF counts failures that emerge only under the full three-update composition, and 3-CFR is computed as 3-CF/Eligible.
h
Residual CFR ↓
Utility ↑
Runtime Check Rate ↓
Time Cost (s/task) ↓
2
4/157 (2.55%)
60.5%
30.1%
30.1
3
2/157 (1.27%)
59.2%
33.6%
32.5
4
2/157 (1.27%)
58.0%
47.6%
49.7
Appendix
Table 6: Sensitivity to the local interaction radius h on AgentDojo.