ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
Authors: Jingnan Zheng, Dongcheng Zhang, Yi Zhang, Ming Zhang, Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He, +5 more
Organizations: Shanghai AI Laboratory · National University of Singapore · Shanghai Jiao Tong University · University of Science and Technology of China · Tsinghua University
Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R2AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.
Figures & tables
Figure 2 : The three-stage ReSI framework. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model Mt , develops training recipes through automated research, and adopts a validated model update Mt+1 as the target for the next round. Stage 1 collects successful attacks against Mt (Sec. 2.1 ). Stage 2 screens attack sources and alignment methods through low-budget pilots and conditional scaling probes, shortlisting up to K recipes (Sec. 2.3 ). Stage 3 conducts full-scale training and feedback-guided refinement, then selects the best eligible checkpoint after all trials conclude (Sec. 2.4 ).
Figure 3 : Safety performance and general capabilities across four model families. Each radar plot compares a backbone model with ReSI and the available safety-training baselines, A3 and MAGIC. The four safety axes (WildJailbreak, HarmBench, H-CoT and X-Teaming) report 100%−ASR . XS-Test reports the response rate on benign requests, with higher scores indicating less over-refusal. GPQA-Diamond and MMLU-Pro report accuracy, and IFEval reports loose instruction-level accuracy. All axes range from 0 to 100%, with higher values indicating better performance. Evaluation protocols and numerical results are provided in Section 3 , Tables 1 and 2 .
Model
Method
WJB ↓
HB ↓
HC ↓
X-Teaming ↓
XS ↑
Qwen
Backbone
11.05
3.75
72.00
80.50
96.00
A3
4.55
0.50
40.00
64.15
91.60
MAGIC
0.40
0.25
32.00
16.98
92.00
ReSI
0.45
0.00
4.00
16.98
98.80
Ministral
Backbone
85.15
53.50
98.00
100.00
96.00
A3
20.35
6.50
76.00
95.60
90.40
Table 1: Safety and benign-compliance results (%). WJB, HB, and HC denote ASRs on WildJailbreak, HarmBench, and H-CoT, respectively; XS denotes the benign-response rate on XS-Test. Bold and underlined values mark the best and second-best distinct scores for each model, with ties formatted identically. MAGIC is omitted for Safin because its current training framework lacks reinforcement-learning support (Appendix B.2 ).
Model
Method
GPQA-Diamond
MMLU-Pro
IFEval prompt
IFEval instruction
Qwen
Backbone
83.84
85.70
92.42
94.48
A3
86.36
85.00
93.16
94.96
MAGIC
86.87
86.00
92.05
94.72
ReSI
86.87
84.90
90.20
93.17
Ministral
Backbone
71.21
73.40
80.59
86.45
A3
66.67
68.70
73.20
81.77
Table 2: General-capability scores (%). Higher values indicate better performance. IFEval reports loose prompt-level and instruction-level accuracy. Bold and underlined values mark the best and second-best distinct scores for each model, with ties formatted identically.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Control
Default
Scope
Initial sample count n0
512
Successful attacks per eligible source.
Critic SFT / A3 pilot budget
128
Optimizer updates per trial.
Safety GRPO pilot budget
8
Rollout–training iterations per trial.
Initial mixture λ(0)
(1,0,0)
Current safety / safety replay / general tasks.
Scaling-probe count nscale
2,048
Target count, subject to verified data supply.
Scaling probes per round
1
Maximum additional scaling trials.
Appendix
Table 3: Configurable pilot-screening defaults. Attack counts refer to verified successful records per source before method-specific data construction.
Skill
Data input
Operation
S1. Source composition
Verified attacks from promising sources.
Add a source, deduplicate, and rebuild supervision. Start with at most two sources ; keep the method, learning rate, and update budget fixed.
S2. Benign pairing
Harmful prompts and generated benign counterparts; exclude XS-Test.
Generate and verify one benign counterpart per unique harmful prompt . Reuse verified pairs within the branch. Initially raise benign sampling weight from 1 to 2 or 3 .
S3. Utility preservation
DOLCI or custom prompt–answer pairs (preparation below).
For SFT, start with 15% or 20% general-task examples while preserving the current-safety/replay ratio. Alternatively, scale learning rate or updates by 0.5 . Change one dimension per trial; GRPO applies the latter adjustments without general-task replay.
S4. Post-SFT safety GRPO
The selected recipe’s candidate prompt pool; exclude general tasks.
Retain its attack sources, benign pairs, and safety replay. Screen prompts and continue safety GRPO from the SFT checkpoint; initialization and screening are detailed below.
Appendix
Table 4: Data and operations for refinement skills. Inputs come from training partitions. Numerical choices are configurable starting values.
Condition
Threshold
Reference
Safety gain
≥5 pp
ASR reduction from Mt on the round’s safety metric.
Instruction-following decline
≤6 pp
Relative to M0 .
Benign-compliance decline
≤6 pp
Relative to M0 .
WJB ASR increase
≤2 pp
Relative to Mt when H-CoT determines safety gain.
Appendix
Table 5: S4 recovery thresholds. All conditions must hold. These configurable thresholds permit further training, not final acceptance; pp denotes percentage points.
Dataset
Evaluation role
Examples
WildJailbreak
In-distribution safety
2,000
HarmBench
Retention of existing safety defenses
400
H-CoT
Safety against challenging attacks
50
X-Teaming
Safety against adaptive attacks
159
XS-Test
Benign compliance
250
GPQA-Diamond
Reasoning
198
Appendix
Table 6: Final evaluation panels. Counts refer to evaluated requests, behaviours, or prompts.
Measure
Validation examples
Reference / requirement
Safety gain gt
128 held-out WildJailbreak requests
Mt ; positive gain on the fixed panel
Low-ASR configuration
50 H-CoT requests if pre-training WJB ASR <5%
H-CoT gain >0 ; WJB ASR increase ≤2 percentage points from Mt
Table 10: Ministral recipes by outer round and training stage. Round-0 GRPO starts from the same round’s SFT model; round 1 continues directly from the retained round-0 model.
Table 11: Training recipes for the selected DeepSeek and Safin checkpoints.
Attack-data source
WJB ASR ↓
H-CoT ASR ↓
Benign compliance ↑
IF ↑
A3
0
30
99.6
92
JailbreakSkill
0
16
99.6
90
MAGIC
0.78
30
99.2
92
Appendix
Table 12: Pilot recipe comparison (%). All candidates adopt critic SFT. IF denotes IFEval loose prompt accuracy throughout this case; bold marks the lowest ASR, including ties.
Checkpoint
WJB ASR ↓
H-CoT ASR ↓
Benign compliance ↑
IF ↑
Decision
Expanded SFT
2.34
26
98
92
Refine
GRPO step 20
0
4
98
96
Select
Appendix
Table 13: Full-scale feedback and final selection (%). Retention checks use 100 benign XS-Test requests and 50 IFEval prompts, with reference scores of 100% and 98% on M0 . Each may decline by at most three percentage points.