cs.AIOct 8, 2026

Recursive Self-Improvement through Multi-Agent Self-Supervision

Authors: Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Somayeh Sojoudi, Matei Zaharia, Yujin Tang

Organizations: UC Berkeley · SakanaAI · Sakana AI

Abstract

Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement

    Sep 28, 2026Yaxin Du, Xiyuan Yang, Zhifan Zhou +10Language Model Post-TrainingSelf-Improving Agents

  2. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

    Sep 21, 2026Peng Xia, Rujun Han, Zifeng Wang +11OOD GeneralizationAgent Harness Optimization

  3. Recursive Harness Self-Improvement

    Jul 17, 2026Hyunin Lee, Jinglue Xu, Jeffrey Seely +3Continual Learning for LLM AgentsAgent Harness Optimization