Learning from Research: Toward Lifelong Agent Harness Evolution
Organizations: University of California, Santa Barbara · Microsoft
Abstract
Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.
Figures & tables
| Qwen3.5-27B | GPT-5.4-mini | |||
| Variant | TGC | SGC | TGC | SGC |
| ScholarEvolve | 81.4 1.19 | 69.0 4.10 | 72.4 1.83 | 55.4 3.55 |
| topic-guided selection | 80.0 0.91 | 66.7 4.12 | 69.8 0.91 | 50.0 4.72 |
| module-wise mutation | 78.6 1.79 | 64.9 4.49 | 69.3 1.24 | 48.8 3.72 |
| research guidance | 76.8 1.19 | 62.5 3.57 | 67.5 2.44 | 45.8 2.70 |
| Qwen3.5-27B | GPT-5.4-mini | |||
| Configuration | TGC | SGC | TGC | SGC |
| Initial harness | 67.8 1.01 | 43.9 3.04 | 66.1 4.43 | 47.3 9.12 |
| Tool only | 72.5 5.36 | 50.9 13.25 | N/A | |
| Context only | 74.9 2.68 | 57.9 5.26 | N/A | |
| Skills only | 73.7 7.02 | 47.4 13.93 | 69.0 3.65 | 52.6 5.25 |
| Memory only | 81.3 1.01 | 59.6 3.04 | 72.5 2.02 | 45.6 3.06 |
| Reference APIs | |||
| Backbone | ( =168) | 10–12 ( =147) | ( =102) |
| Qwen3.5-27B | 13.1 | 18.4 | 9.5 |
| GPT-5.4-mini | 5.4 | 15.9 | 7.2 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Metric | Method | ||||
| TGC | ScholarEvolve | ||||
| Meta Harness | |||||
| SGC | ScholarEvolve | ||||
| Meta Harness |
| Normal | Challenge | Recovered | |||
| Subtype | Initial | Evolved | Initial | Evolved | (pooled) |
| Retrieval & grounding | |||||
| Incomplete retrieval | 12 | 8 | 65 | 62 | 0 |
| Entity confusion | 7 | 11 | 8 | 12 | 0 |
| Unsupported inference | 1 | 1 | 3 | 5 | 0 |
| Goal & reasoning | |||||
| Module | Topic | Source paper | 90% interval | W/L | ||
| Tool interface | Action Space Control | Budget-constrained tool learning ( Zheng et al., 2024 ) | +2.9 | [ 3.5, +9.4] | 15/11 | |
| Abstract Strategy Actions | CheMatAgent ( Wu et al., 2025 ) | +1.2 | [ 4.1, +6.4] | 12/10 | ||
| Structured API Actions | ToolACE-R ( Zeng et al., 2026b ) | +4.7 | [ 0.6, +9.9] | 14/8 | ||
| Constructed Composite Actions | RefTool ( Liu et al., 2026b ) | +1.8 | [ 4.7, +8.2] | 15/10 | ||
| Context | Lossy Context Rewriting | Compression without MLPs ( Honig et al., 2025 ) | 45.6 | [ 53.8, 36.8] | 1/40 | |
| Independent Unit Scoring | ICPC ( Yu & Liu, 2025 ) | 43.3 | [ 52.0, 34.5] | 2/40 |
| AppWorld | -Bench Telecom | |||
| Module | Qwen3.5-27B | GPT-5.4-mini | Qwen3.5-27B | GPT-5.4-mini |
| Tool interface | ToolACE-R ( Zeng et al., 2026b ) Structured API Actions | – | Lower Privileges Suffice ( Yang et al., 2026b ) Action-Centric Post-Training | Interactive Semantic Parsing ( Yao et al., 2019 ) Action Validation and Repair |
| Context | Prompt compression limits ( Nagle et al., 2024 ) Budget and Scope Routing | – | Diable ( Lesci et al., 2023 ) Structured State Memory | Measure Before You Manage ( Chen et al., 2026 ) Context Buffer Management |
| Skills | Skill-as-Pseudocode ( Li et al., 2026 ) Skill Compilation | Skill Drift Is Contract Violation ( Fan et al., 2026 ) Skill Library Maintenance | Procedural KG Extraction ( Carriero et al., 2024 ) Skill Induction | Procedural Knowledge at Scale ( Wu et al., 2026a ) Procedural Memory and Retrieval |
| Memories | Oracle Agent Memory ( Alake et al., 2026 ) Structured Memory Architecture | CoEvo-Mem ( Ye et al., 2026 ) Adaptive Retrieval Control | CAST ( Ma et al., 2026 ) Structured Relational Memory | MemRL ( Zhang et al., 2026d ) Adaptive Retrieval and Retention |
| Workflow | TMAS ( Wu et al., 2026b ) Multi-Agent Deliberation | Plan-and-Solve ( Wang et al., 2023 ) Plan-and-Execute | Textual-to-Visual Self-Verification ( Xu et al., 2025 ) Hierarchical Decomposition | Same-Model Self-Verification ( Phalod, 2026 ) Adaptive Routing |