cs.AIOct 7, 2026

RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement

Authors: Renxiong Wang, Darvin Yi, Abril Herrlein, Anas Mahmoud, Advait Gosai, Lisiman Hua, MohammadHossein Rezaei, Xingang Guo, +11 more

Organizations: Scale AI · University of California, Santa Cruz · University of North Carolina at Chapel Hill

Abstract

Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to produce tasks with automated evaluators; each paper's method is independently reimplemented to establish a baseline score. We present 210 environments across 18 fields, including 90 reviewed by independent human domain experts. Both experts and agent judges give high ratings to the potential for improving the provided starting solutions and the evaluators' ability to distinguish solution quality, whereas experts are more critical of shortcut resistance, faithfulness to the source paper, and whether a single idea can exhaust a task. To validate their use for repeated improvement, we evaluate four models over 3 successive attempts on 120 environments, with each attempt inheriting prior code and notes while model weights remain fixed. At least one model improves after the first attempt in 84% of environments. Models also outperform the reproduced paper methods in 68 of the 120 environments, demonstrating room for gains beyond these baselines. Transcript analysis identifies work beyond parameter tuning in 95% of these successful attempts. Analysis of the resulting trajectories shows that models scoring lower on these tasks explore less, more often accept gains smaller than the reported standard error, and rely more heavily on tuning to the development set. RSI-Forge provides a scalable approach to constructing research environments for training and evaluating self-improving agents.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RSIGym: A Flexible Environment for Recursive Self-Improvement

    Oct 7, 2026Fanqing Meng, Lingxiao Du, Haocheng Lu +6LLM Agent Self-ImprovementRecursive Self-Improvement

  2. RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

    Sep 14, 2026Sibo Zhu, Shicheng Fan, Xinyue Wang +3Multi-Agent LLM SystemsLLM Agent Self-Improvement

  3. RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement

    Sep 28, 2026Yaxin Du, Xiyuan Yang, Zhifan Zhou +10Language Model Post-TrainingSelf-Improving Agents