EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning
Organizations: Fudan University · Renmin University of China
Abstract
Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework that bridges evolution and internalization. EvoIn first analyzes agent execution traces to evolve and validate new decision-making procedures by temporarily instantiating them in the harness. The validated procedures guide the agent to generate improved reasoning traces. These traces are then rewritten into self-contained reasoning traces, removing explicit references to harness instructions while expressing the induced decision logic as the model's own reasoning. Finally, EvoIn fine-tunes the model on the rewritten traces, internalizing these procedures so that the improved decision-making persists without the evolved harness at inference time. We evaluate EvoIn on diverse benchmarks and find that it consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain. Results further show that the internalized decision procedures generalize to unseen tasks. Case studies show that agents can learn to decide how to solve a task before solving it, for example by checking a document's length to choose between reading it in full and searching it. EvoIn is also broadly applicable, showing consistent improvements on another model family.
Figures & tables
| Pass | Score | ||||||
| Task group | Base | EvoIn | Base | EvoIn | |||
| Multi-document key retrieval | 300 | 57.00 | 63.33 | 6.33 | 57.00 | 63.33 | 6.33 |
| Evidence-grounded QA | 400 | 52.25 | 61.50 | 9.25 | 55.03 | 64.93 | 9.90 |
| Structured-data reasoning | 300 | 44.00 | 61.33 | 17.33 | 44.00 | 61.33 | 17.33 |
| Log and dialogue tracking | 300 | 53.00 | 75.00 | 22.00 | 55.14 | 76.78 | 21.64 |
| In-context learning | 200 | 26.00 | 39.00 | 13.00 | 26.00 | 39.00 | 13.00 |
| Pass | Score | ||||||
| Benchmark | Base | EvoIn | Base | EvoIn | |||
| AA-LCR | 100 | 27.00 | 37.00 | 10.00 | 27.00 | 37.00 | 10.00 |
| BrowseComp-LongContext | 295 | 12.54 | 13.56 | 1.02 | 12.54 | 13.56 | 1.02 |
| LongBench v2 | 503 | 33.20 | 44.73 | 11.53 | 33.20 | 44.73 | 11.53 |
| MRCR | 800 | 2.00 | 8.88 | 6.88 | 32.55 | 76.31 | 43.76 |
| Oolong | 300 | 43.33 | 60.33 | 17.00 | 44.58 | 61.42 | 16.84 |
| ID | OOD | ||||
| Method | Corpus size | Pass | Score | Pass | Score |
| Base | 0 | 35.96 | 44.58 | 18.99 | 31.02 |
| Seed-harness SFT | 12,035 | 37.83 | 47.25 | 18.36 | 43.16 |
| EvoIn w/o Tailor | 12,035 | 45.00 | 54.89 | 19.67 | 45.70 |
| EvoIn | 12,035 | 46.83 | 56.40 | 28.20 | 54.49 |
| Qwen rollout | 1,752 | 46.52 | 56.30 | 22.04 | 49.36 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Task |
| Multi-document key retrieval | |
| Multi-Doc Key Lookup | Locate randomly generated keys in synthetic multi-document inputs, then answer the questions attached to them or combine their answers. |
| Needle QA | Find the question that follows a given key in a specified document amid long distractor text, then answer it. |
| Cross-Doc Key Aggregation | Evaluate keyed expressions across all documents and aggregate the results, e.g., by sum, maximum, or conditional operations. |
| Evidence-grounded QA | |
| Exam Reading Comprehension | Answer multiple-choice, translation, and short-answer questions on Chinese college-entrance-exam reading passages. |
| Pass | Score | |||||
| Category | Base | EvoIn | Base | EvoIn | ||
| Multi-document key retrieval | ||||||
| Multi-Doc Key Lookup | 77.00 | 83.00 | 6.00 | 77.00 | 83.00 | 6.00 |
| Needle QA | 61.00 | 55.00 | 6.00 | 61.00 | 55.00 | 6.00 |
| Cross-Doc Key Aggregation | 33.00 | 52.00 | 19.00 | 33.00 | 52.00 | 19.00 |
| Evidence-grounded QA | ||||||
| Setting | Value |
| Fine-tuning | Full-parameter SFT |
| Framework | Megatron-Bridge ( Shoeybi et al., 2019 ) |
| Epochs | 3 |
| Sequence length | 24,576 tokens |
| Global / micro batch size | 8 / 1 |
| Tensor / expert parallelism | 8 / 8 |
| Setting | Evaluated model | Qwen3.5-35B-A3B judge | GPT-OSS-120B judge |
| Context limit (tokens) | 180,224 | – | – |
| Max generated tokens | 16,000 | 8,192 | 8,192 |
| Temperature | 1.0 | 0 | 0.7 |
| Top- | 1.0 | 1.0 | 0.95 |
| Top- | 50 | – | 50 |
| Repetition penalty | 1.05 | – | 1.05 |
| Qwen3.5-35B-A3B | Gemma | ||||||||
| Category | Base | EvoIn | Seed- harness | w/o Tailor | All- Qwen | Qwen rollout | GLM-5.3 rollout | Base | EvoIn |
| Multi-document key retrieval | |||||||||
| Multi-Doc Key Lookup | 77.00 | 83.00 | 74.00 | 86.00 | 84.00 | 87.00 | 89.00 | 75.00 | 86.00 |
| Needle QA | 61.00 | 55.00 | 62.00 | 62.00 | 66.00 | 73.00 | 79.00 | 59.00 | 76.00 |
| Cross-Doc Key Aggregation | 33.00 | 52.00 | 35.00 | 55.00 | 43.00 | 47.00 | 70.00 | 46.00 | 53.00 |
| Evidence-grounded QA | |||||||||
| Qwen3.5-35B-A3B | Gemma | |||||||||
| Benchmark | Base | EvoIn | Seed- harness | w/o Tailor | All- Qwen | Qwen rollout | GLM-5.3 rollout | Base | EvoIn | |
| Pass | ||||||||||
| AA-LCR | 100 | 27.00 | 37.00 | 30.00 | 36.00 | 34.00 | 37.00 | 49.00 | 58.00 | 64.00 |
| BrowseComp-LongContext | 295 | 12.54 | 13.56 | 11.19 | 12.20 | 15.25 | 12.88 | 19.32 | 22.03 | 32.88 |
| LongBench v2 | 503 | 33.20 | 44.73 | 25.25 | 21.87 | 36.58 | 36.98 | 36.78 | 58.05 | 61.43 |
| MRCR | 800 | 2.00 | 8.88 | 6.88 | 9.50 | 3.00 | 5.75 | 15.00 | 50.38 | 43.12 |