cs.AISep 29, 2026

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

Authors: Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, +6 more

Organizations: Beijing Academy of Artificial Intelligence (BAAI)

Abstract

We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.

Figures & tables

Explore similar work

CardsList
  1. Retrospective Progress-Aware Self-Refinement for LLM Agent Training

    Jun 12, 2026Xinbei Ma, Congmin Zheng, Jiyang Qiu +10Self-Improving AgentsLarge Language Model Agents

  2. AEL: Agent Evolving Learning for Open-Ended Environments

    Apr 23, 2026Wujiang Xu, Jiaojiao Han, Minghao Guo +4Self-Evolving AgentsRecursive Self-Improvement

  3. SelfSearch: Reward-Free Search for Self-Improving Agents

    Sep 29, 2026Jungwoo Yang, Injin Kong, Yohan JoSelf-Improving AgentsSelf-Evolving Agents