cs.CROct 8, 2026

ReSI: Recursive Safety Improvement toward Resistant and Resilient AI

Authors: Jingnan Zheng, Dongcheng Zhang, Yi Zhang, Ming Zhang, Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He, +5 more

Organizations: Shanghai AI Laboratory · National University of Singapore · Shanghai Jiao Tong University · University of Science and Technology of China · Tsinghua University

Abstract

Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R2^2AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement

    Sep 28, 2026Yaxin Du, Xiyuan Yang, Zhifan Zhou +10Language Model Post-TrainingSelf-Improving Agents

  2. Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

    May 9, 2026Dongcheng Zhang, Yi Zhang, Yuxin Chen +3LLM Self-CorrectionRL for Language Model Reasoning