Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework that bridges evolution and internalization. EvoIn first analyzes agent execution traces to evolve and validate new decision-making procedures by temporarily instantiating them in the harness. The validated procedures guide the agent to generate improved reasoning traces. These traces are then rewritten into self-contained reasoning traces, removing explicit references to harness instructions while expressing the induced decision logic as the model's own reasoning. Finally, EvoIn fine-tunes the model on the rewritten traces, internalizing these procedures so that the improved decision-making persists without the evolved harness at inference time. We evaluate EvoIn on diverse benchmarks and find that it consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain. Results further show that the internalized decision procedures generalize to unseen tasks. Case studies show that agents can learn to decide how to solve a task before solving it, for example by checking a document's length to choose between reading it in full and searching it. EvoIn is also broadly applicable, showing consistent improvements on another model family.
Figures & tables
Figure 1: Four paradigms for improving language agents. (A) Model fine-tuning updates the model while keeping the harness fixed. (B) Harness evolution updates the harness while keeping the model fixed. (C) Model–harness co-evolution updates both jointly. (D) EvoIn first evolves the harness to discover improved decision procedures, then internalizes the reusable procedures into the model and returns to the original harness. Flame icons and superscript ∗ indicate updated components.
Figure 2: We illustrate EvoIn with a multi-document reasoning case. EvoIn refines decision procedures through iterative proposal and evaluation, then samples and rewrites trajectories for fine-tuning. Gray highlights mark procedure updates in the evolved harness. Red highlights show explicit reliance on instructions from the evolved procedures, while green highlights show the same decision logic reformulated as the agent’s own reasoning.
Pass
Score
Task group
N
Base
EvoIn
Δ
Base
EvoIn
Δ
Multi-document key retrieval
300
57.00
63.33
+ 6.33
57.00
63.33
+ 6.33
Evidence-grounded QA
400
52.25
61.50
+ 9.25
55.03
64.93
+ 9.90
Structured-data reasoning
300
44.00
61.33
+ 17.33
44.00
61.33
+ 17.33
Log and dialogue tracking
300
53.00
75.00
+ 22.00
55.14
76.78
+ 21.64
In-context learning
200
26.00
39.00
+ 13.00
26.00
39.00
+ 13.00
Table 1: ID results by task group (%). Base is Qwen3.5-35B-A3B, and EvoIn is the same model after EvoIn fine-tuning (epoch 3). Each group value is the mean over its categories, each with 100 evolution-test examples; Appendix A.1 lists the categories and their results. Δ is EvoIn minus Base, and Pass equals Score for groups whose categories all have binary scores.
Pass
Score
Benchmark
N
Base
EvoIn
Δ
Base
EvoIn
Δ
AA-LCR
100
27.00
37.00
+ 10.00
27.00
37.00
+ 10.00
BrowseComp-LongContext
295
12.54
13.56
+ 1.02
12.54
13.56
+ 1.02
LongBench v2
503
33.20
44.73
+ 11.53
33.20
44.73
+ 11.53
MRCR
800
2.00
8.88
+ 6.88
32.55
76.31
+ 43.76
Oolong
300
43.33
60.33
+ 17.00
44.58
61.42
+ 16.84
Table 2: OOD results by benchmark (%). Models and Δ follow Table 1 .
ID
OOD
Method
Corpus size
Pass
Score
Pass
Score
Base
0
35.96
44.58
18.99
31.02
Seed-harness SFT
12,035
37.83
47.25
18.36
43.16
EvoIn w/o Tailor
12,035
45.00
54.89
19.67
45.70
EvoIn
12,035
46.83
56.40
28.20
54.49
Qwen rollout
1,752
46.52
56.30
22.04
49.36
Table 3: Ablation study (%). All rows use the same Qwen student, SFT schedule, seed harness, and evaluation protocol as Tables 1 and 2 . Corpus size counts SFT demonstrations. The top block separately ablates harness evolution and rewriting. The bottom block compares EvoIn using equally sized Qwen and GLM-5.3 rollout subsets from the same tasks. Bold marks each block’s best result.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Task
Multi-document key retrieval
Multi-Doc Key Lookup
Locate randomly generated keys in synthetic multi-document inputs, then answer the questions attached to them or combine their answers.
Needle QA
Find the question that follows a given key in a specified document amid long distractor text, then answer it.
Cross-Doc Key Aggregation
Evaluate keyed expressions across all documents and aggregate the results, e.g., by sum, maximum, or conditional operations.
Evidence-grounded QA
Exam Reading Comprehension
Answer multiple-choice, translation, and short-answer questions on Chinese college-entrance-exam reading passages.
Appendix
Table 4: ID task categories , grouped by the task groups in Table 1 .
Pass
Score
Category
Base
EvoIn
Δ
Base
EvoIn
Δ
Multi-document key retrieval
Multi-Doc Key Lookup
77.00
83.00
+ 6.00
77.00
83.00
+ 6.00
Needle QA
61.00
55.00
− 6.00
61.00
55.00
− 6.00
Cross-Doc Key Aggregation
33.00
52.00
+ 19.00
33.00
52.00
+ 19.00
Evidence-grounded QA
Appendix
Table 5: ID results by category (percent). Models and Δ follow Table 1 , each category has 100 evolution-test examples.
Setting
Value
Fine-tuning
Full-parameter SFT
Framework
Megatron-Bridge ( Shoeybi et al., 2019 )
Epochs
3
Sequence length
24,576 tokens
Global / micro batch size
8 / 1
Tensor / expert parallelism
8 / 8
Appendix
Table 6: SFT configuration for the main model and all Qwen ablations.
Setting
Evaluated model
Qwen3.5-35B-A3B judge
GPT-OSS-120B judge
Context limit (tokens)
180,224
–
–
Max generated tokens
16,000
8,192
8,192
Temperature
1.0
0
0.7
Top- p
1.0
1.0
0.95
Top- k
50
–
50
Repetition penalty
1.05
–
1.05
Appendix
Table 7: Decoding settings in the Qwen evaluations. The GPT-OSS-120B column lists the judge settings for the OOD benchmarks. A dash marks a setting that is not set explicitly.
Qwen3.5-35B-A3B
Gemma
Category
Base
EvoIn
Seed- harness
w/o Tailor
All- Qwen
Qwen rollout
GLM-5.3 rollout
Base
EvoIn
Multi-document key retrieval
Multi-Doc Key Lookup
77.00
83.00
74.00
86.00
84.00
87.00
89.00
75.00
86.00
Needle QA
61.00
55.00
62.00
62.00
66.00
73.00
79.00
59.00
76.00
Cross-Doc Key Aggregation
33.00
52.00
35.00
55.00
43.00
47.00
70.00
46.00
53.00
Evidence-grounded QA
Appendix
Table 8: Per-category ID Score of all runs (percent). The ablation columns correspond to the questions in Section 5 : Seed-harness SFT removes harness evolution, EvoIn w/o Tailor removes rewriting, All-Qwen uses the target model as both the proposer and the Tailor model, the two rollout columns are the task-paired Qwen and GLM-5.3 arms, and the Gemma columns repeat the pipeline with Gemma-4-31B-it.
Qwen3.5-35B-A3B
Gemma
Benchmark
N
Base
EvoIn
Seed- harness
w/o Tailor
All- Qwen
Qwen rollout
GLM-5.3 rollout
Base
EvoIn
Pass
AA-LCR
100
27.00
37.00
30.00
36.00
34.00
37.00
49.00
58.00
64.00
BrowseComp-LongContext
295
12.54
13.56
11.19
12.20
15.25
12.88
19.32
22.03
32.88
LongBench v2
503
33.20
44.73
25.25
21.87
36.58
36.98
36.78
58.05
61.43
MRCR
800
2.00
8.88
6.88
9.50
3.00
5.75
15.00
50.38
43.12
Appendix
Table 9: Per-benchmark OOD results of all runs (percent). Columns follow Table 8 . The Qwen-rollout and Gemma columns use the same GPT-OSS-120B judge served from a different endpoint, and Gemma also uses its own evaluation protocol.
A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks come from many domains, so both the evolution and the evaluation of a harness should cover a diverse range of tasks. We propose a framework close to recursive self-improvement: the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it. Each evolution batch draws tasks from five benchmarks in different domains. To measure generalization, training and held-out tasks are strictly separated, and we additionally evaluate on five out-of-distribution benchmarks never used during evolution. We frame the evolution process as deep-learning training with two stages, multi-task pretraining and continual training. Starting from a 49-line seed harness, the harness obtained at the end of the first stage improves the average score by 4.48 points on the in-distribution benchmarks and by 12.64 points on the out-of-distribution benchmarks, surpassing Codex on the former and matching it on the latter. In the second stage, continued evolution on Claw-Eval, one of the out-of-distribution benchmarks, further raises the score on that benchmark from 66.17 to 68.06, exceeding Codex. We also provide an in-depth analysis of the mechanisms that emerged during evolution, including output truncation, history compaction, and independent review.
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
Lisheng Huang, Chen Yang, Hao Zhou +6
1Gaoling School of Artificial Intelligence, Renmin University of China · 2BOSS Zhipin, Beijing, China
Large language models (LLMs) exhibit strong reasoning capabilities, yet most LLM-based agents are statically deployed and unable to improve through task interactions. Existing experience-driven methods often rely on memory or heuristics without enhancing the model's ability to learn, treating it as a passive executor and leading to early performance plateaus and limited long-term improvement. To address this issue, we propose MetaEvo, a two-stage framework for continual agent evolution that focuses on improving how the model learns from tasks experience, rather than solely on what it stores. MetaEvo first applies preference-based optimization to enhance the model's ability of principle abstraction, then enables the accumulation and reuse of these principles within a modular agent architecture. Experimental results on diverse reasoning benchmarks demonstrate that MetaEvo consistently outperforms strong baselines, maintains reliable improvement across iterations. These findings validate the effectiveness of meta-optimization in enabling agents to learn from experience and continually enhance their reasoning capabilities.
Bowen Ren, Heyan Huang, Yinghao Li +1
School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China · Beijing Institute of Technology Southeast Academy of Information Technology, Putian, China