We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.
Figures & tables
Figure 1 : Results of AREX-2 on six benchmarks, compared with selected closed and open models.
Figure 2 : Model size versus performance on four benchmarks. AREX-2 (27B) matches or exceeds much larger models.
Figure 3 : Overview of AREX-2. Left: environments are constructed from GitHub repositories and online judges, and are kept only if they pass execution and validation. Middle: an agent improves its solution over many rounds, acquiring operational knowledge on demand, and its trajectories become training data. Right: the model trained on this data improves its solution over multiple rounds at test time.
Model
Size
Frontier-CS
MLE-Lite
Closed-Weight Models
GPT-5.6 Sol
–
76.4
72.7
Claude Opus 4.8
–
74.5
63.6
GPT-5.5
–
72.1
68.2
Gemini-3.1-Pro
–
68.9
–
Qwen3.7-Max
–
61.9
–
Table 1 : Comparison on coding and machine learning engineering benchmarks. Model size refers to the total number of parameters. Frontier-CS results are reported on the 188-task Agent Track. ∗ denotes results reproduced by us. MLE-Lite results are reproduced and reported by OpenMLE. MLE-Lite reports Any Medal (mean over three seeds) on MLE-bench Lite. AREX-2 is evaluated on MLE-Lite with skills in its context ( \Cref sec:case-mle).
Model
Size
BrowseComp
HLE
GAIA
DeepSearchQA
Frontier Models
GPT-5.6 Sol
–
90.4
58.0 *
–
–
GPT-5.6 Terra
–
87.5
–
–
–
GPT-5.6 Luna
–
83.3
–
–
–
Kimi-K3
2.8T
91.2
56.0 *
–
95.0
Claude Fable 5
–
88.0
64.5 *
–
94.2
Table 2 : Comparison on general agentic reasoning and deep research benchmarks. ∗ denotes results reported on the full HLE set; unmarked results use the text-only subset.
Figure 6
Figure 6 : Stage-wise ablation on MLE-bench Lite. Stages M0 to M2 use the base model, and M3 and M4 use AREX-2.
LLM-based agents trained with reinforcement learning optimize step-wise action prediction but lack metacognitive awareness of task progress, inducing a gap that hinders long-horizon scaling. A pilot study reveals that online progress prompting hurts performance while retrospective demonstrations help, yet this capability cannot emerge from outcome-reward training alone. We present RePro, Retrospective Progress-Aware Training, a framework that trains agents to self-generate progress signals via a forward-then-reflect rollout paradigm: the agent executes actions online, then retrospectively reassesses its step-wise progress given the completed trajectory and known outcome. RePro initializes with a Retrospection Warmup that teaches reflection format from minimal external demonstrations, then further trains through RePro-PO with a composite reward that produces self-generated signals without continuous external supervision. Experiments on WebShop, ALFWorld, and Sokoban show that RePro enhances the Qwen family's performance, with up to 12% absolute success rate gains.
Xinbei Ma, Congmin Zheng, Jiyang Qiu +10
1Shanghai Jiao Tong University · 2OPPO Research Institute
LLM agents increasingly operate in open-ended environments spanning hundreds of sequential episodes, yet they remain largely stateless: each task is solved from scratch without converting past experience into better future behavior. The central obstacle is not \emph{what} to remember but \emph{how to use} what has been remembered, including which retrieval policy to apply, how to interpret prior outcomes, and when the current strategy itself must change. We introduce \emph{Agent Evolving Learning} (\ael{}), a two-timescale framework that addresses this obstacle. At the fast timescale, a Thompson Sampling bandit learns which memory retrieval policy to apply at each episode; at the slow timescale, LLM-driven reflection diagnoses failure patterns and injects causal insights into the agent's decision prompt, giving it an interpretive frame for the evidence it retrieves. On a sequential portfolio benchmark (10 sector-diverse tickers, 208 episodes, 5 random seeds), \ael{} achieves a Sharpe ratio of 2.13±0.47, outperforming five published self-improving methods and all non-LLM baselines while maintaining the lowest variance among all LLM-based approaches. A nine-variant ablation reveals a ``less is more'' pattern: memory and reflection together produce a 58% cumulative improvement over the stateless baseline, yet every additional mechanism we test (planner evolution, per-tool selection, cold-start initialization, skill extraction, and three credit assignment methods) \emph{degrades} performance. This demonstrates that the bottleneck in agent self-improvement is \emph{self-diagnosing how to use} experience rather than adding architectural complexity. Code and data: https://github.com/WujiangXu/AEL.
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{$4.03} in search cost, it produces a harness that solves \textbf{82.0}% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.
Jungwoo Yang, Injin Kong, Yohan Jo
Graduate School of Data Science, Seoul National University