Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
Organizations: School of Data Science, Fudan University · Shanghai Key Laboratory of Data Science · College of Computer Science and Artificial Intelligence, Fudan University
Abstract
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce ContextEvo, a framework that learns a context policy from long-horizon trajectories. ContextEvo reconstructs the model-visible context at key decision points, identifies context-related failures, and applies targeted policy updates. Starting from the open-source Pi-agent harness, ContextEvo improves performance across three long-horizon task benchmarks, achieving results comparable to or better than several prominent agent harnesses, including Codex, OpenCode, and OpenClaw. Additional analyses show that fixed or locally evolved context strategies can fall short under long-horizon information pressure, while our methods adapt to the information demands of each environment.
Figures & tables
| Dimension | Role | Common realizations |
|---|---|---|
| Input Assembly | Context entry preparation | Tool-result formatting; observation compression |
| History Maintenance | Historical context management | Compaction; summarization; persistent state |
| Context Orchestration | Context access and sharing across agents and subtasks | Instructions and tools for context access; subagent instructions and feedback design |
| Long-Horizon TB | DeepSWE | BrowseComp-Plus | |||||||
| Harness / policy | Held-in (%) | Held-out (%) | All (%) | Held-in (%) | Held-out (%) | All (%) | Held-in (%) | Held-out (%) | All (%) |
| Harness (DeepSeek-V4-Flash) | |||||||||
| OpenCode | 36.1 | 40.3 | 38.5 | 62.5 | 66.2 | 64.6 | 92.1 | 77.6 | 83.9 |
| Codex | 41.8 | 40.1 | 40.8 | 45.8 | 63.1 | 55.8 | 84.2 | 84.5 | 84.4 |
| OpenClaw | 36.9 | 45.9 | 42.0 | 45.8 | 60.0 | 54.0 | 89.5 | 88.8 | 89.1 |
| Pi-agent (base) | 43.3 | 40.4 | 41.7 | 54.2 | 72.3 | 64.6 | 81.6 | 82.7 | 82.2 |
| Agent | Reward | Tokens (M) |
|---|---|---|
| Baseline (Pi) | 0.417 | 8.57 |
| Skill-only | 0.370 | 9.61 |
| ContextEvo | 0.448 | 7.52 |
| ContextEvo + skill | 0.419 | 9.69 |
| Method | DeepSWE | LHTB | BrowseComp-Plus |
|---|---|---|---|
| Pi-agent base | 0.133 | 0.303 | 0.379 |
| Static context management | |||
| ReSum | 0.062 | 0.266 | 0.385 |
| ACON | 0.124 | 0.262 | 0.374 |
| Dynamic context management | |||
| TACO | 0.115 | 0.281 | 0.374 |
| Variant | Held-in | Held-out | All | vs. Full |
|---|---|---|---|---|
| Native Pi-agent | 43.30% | 40.39% | 41.66% | pp |
| Full ContextEvo | 46.95% | 43.12% | 44.78% | – |
| w/o Paginated Trajectory Review | 41.01% | 26.19% | 32.64% | pp |
| w/o Cross-Case Aggregation | 40.46% | 44.05% | 42.49% | pp |
| w/o Counterfactual Replay | 48.57% | 35.29% | 41.06% | pp |
| w/o Context-Policy Update | 32.57% | 42.78% | 38.34% | pp |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Scenario | All | Held-in | Held-out |
|---|---|---|---|---|
| LHTB | terminal workflows | 46 | 20 | 26 |
| DeepSWE | repository software engineering | 113 | 48 | 65 |
| BrowseComp-Plus | multi-hop deep research | 174 | 76 | 98 |
| Benchmark | Trajectories | Tool calls | Agent turns | Estimated trace tokens |
|---|---|---|---|---|
| LHTB | 44 | |||
| DeepSWE | 112 | |||
| BrowseComp-Plus | 174 |
| Selection stage | Criterion |
|---|---|
| Reference difficulty | o3 and GPT-5 reference answers are both judged incorrect |
| Long-horizon filter | o3 issues at least 30 searches |
| Resulting subset | 174 questions from the 830-question release |
| Evolution split | seed 15; 76 held-in / 98 held-out |
| Setting | Configuration |
|---|---|
| DeepSeek-V4-Flash-0731 | openai/deepseek-v4-flash ; reasoning enabled; context 262,144; max output 32,768 |
| GPT-5.6-Luna | openai/gpt-5.6-luna ; reasoning enabled; context 262,144; max output 32,768 |
| Generation parameters | Temperature and top- use the provider defaults; no additional sampling override |
| Harness | Execution configuration |
|---|---|
| OpenCode | v1.15.13-20260715; OpenAI-compatible protocol; context/output 262,144/32,768; task timeout 5,400 s; 46/32/64 concurrent trials |
| Pi-agent | v0.80.6; openai-responses ; context/output 262,144/32,768; task timeout 5,400 s; 24/32/64 concurrent trials |
| Codex | Frozen comparison release; openai-chat ; task timeout 5,400 s; 46/32/64 concurrent trials |
| OpenClaw | v2026.6.1; openai-completions ; context/output 262,144/32,768; task timeout 5,400 s; 46/32/64 concurrent trials |
| Policy interface | Lifecycle intervention point | Current realization | Examples of possible policy modifications | |
|---|---|---|---|---|
| 1 | Input Assembly | Observation entry | observation_ [-1pt] projector | Change whether a newly received observation is retained in its original form or represented by a shorter, source-linked form. |
| 2 | History Maintenance | Compaction planning | compaction_ [-1pt] planner | Change when historical material becomes eligible for transformation and which complete dependency groups may be selected. |
| 3 | History Maintenance | Historical representation update | summary_writer | Change which facts, constraints, validation results, failed approaches, and source references are preserved in a compact representation. |
| 4 | History Maintenance | Evidence memory update | memory_writer | Change how durable evidence records are added, replaced, or removed and how their provenance is maintained. |
| 5 | History Maintenance | Evidence recovery | memory_ [-1pt] retriever | Change when prior evidence is requested and which narrowly relevant records are recovered. |
| 6 | Input Assembly | Request assembly | context_ [-1pt] assembler | Change how active history, summaries, recovered evidence, and inherited context are selected and ordered. |
| Harness | Evolved | Native baseline | Full-task change |
|---|---|---|---|
| OpenCode | 71/174 (40.80%) | 61/174 (35.06%) | +10 (+5.74 pp) |
| OpenClaw | 79/174 (45.40%) | 33/174 (18.97%) | +46 (+26.44 pp) |
| Transfer setting | Target | DeepSeek | Luna |
|---|---|---|---|
| Same environment, cross-model | LHTB | pp | pp |
| Cross-environment, frozen policy | DeepSWE | pp | pp |
| Cross-environment, frozen policy | BCP | pp | pp |
| Control surface | Evidence | Role |
|---|---|---|
| Observation rendering | Observed in LHTB; related changes elsewhere | Select actionable context |
| History maintenance | Observed across long trajectories | Compact history while protecting active work |
| Durable evidence state | Present in inspected realizations | Preserve facts, open items, and recovery paths |
| Grounding and verification guidance | Loaded in LHTB and DeepSWE | Guide context use and verification |
| Addressable recovery | Registered; reward effect not isolated | Recover displaced evidence by reference |
| Bounded delegation | Exercised in four BrowseComp-Plus calls | Isolate a branch and return checkable evidence |