Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification
Organizations: University of Maryland, Baltimore County · Emergence AI
Abstract
Recursive self-improvement (RSI) enables agents to iteratively optimize their workflows via execution feedback. However, standard RSI typically operates as a first-order optimizer: it repeatedly patches surface-level parameters in response to immediate failure symptoms, often leading to trial-and-error thrashing without resolving underlying mechanisms. To address this limitation, we introduce SO-RSI, a framework that elevates workflow optimization to a second-order diagnostic inquiry, investigating why failures occur before committing to structural interventions. SO-RSI passively monitors execution traces for three structural anomalies (recurrence, opposing edits, and expectation mismatch) to trigger targeted mechanism investigations. By executing lightweight diagnostic probes and maintaining persistent inquiry memory across RSI rounds, SO-RSI accumulates causal evidence to guide systematic workflow edits rather than parameter patches. Across Lean 4 proof generation and Verus-based verifiable code generation, SO-RSI improves final held-out pass rates over Naive RSI by 21.8 and 25.8 percentage points under matched 24-hour search budgets. Behavioral analyses further confirm that SO-RSI substantially suppresses failure recurrence and eliminates unproductive zero-progress optimization loops.
Figures & tables
| Inquiry category | Attempts | Share |
|---|---|---|
| First-Order: Direct repair (F) | 86 | 64.2% |
| Second-Order Behavior (Permissive Grouping) | 48 | 35.8% |
| Mechanism hypothesis, untested ( H ) | 29 | 21.6% |
| Mechanism investigation ( I ) | 8 | 6.0% |
| Evidence-guided intervention ( G ) | 7 | 5.2% |
| Insufficient information ( U ) | 4 | 3.0% |
| Inquiry group | Evaluated edits | Mean gain (pp) | Performance regressions |
|---|---|---|---|
| Direct repair (F) | 51 | +8.7 | 20/51 (39.2%) |
| Untested hypothesis (H) | 17 | +11.2 | 4/17 ( 23.5% ) |
| Diagnostic inquiry (I/G) | 11 | +14.3 | 3/11 (27.3%) |
| Lean proving | Verus VCG | |||
|---|---|---|---|---|
| Method | Pass | Pass | ||
| Initial workflow | 29.4 | 0.0 | 14.8 | 0.0 |
| Naive RSI | +20.0 | +18.6 | ||
| SO-RSI | +41.8 | +44.4 | ||
| Variant | Lean pass | VCG pass |
|---|---|---|
| SO-RSI | ||
| w/o executable probes | ||
| w/o persistent investigation memory |
| Investigation | G edits | Persistence | ||||
| Method | Tested (%) | Revised (%) | Accept (%) | Pass | Rec. | Loops |
| Lean proving | ||||||
| Naive RSI | 14.8 | 0.9 | 54.8 | 4.4 | -7.3 | 18.7 |
| SO-RSI | 46.0 | 22.6 | 67.1 | 9.8 | -31.4 | 7.2 |
| w/o probes | — | — | 56.5 | 6.1 | -12.8 | 13.9 |
| w/o memory | 41.6 | 15.3 | 59.7 | 7.0 | -19.6 | 11.8 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Criterion | Evidence required | Insufficient evidence |
|---|---|---|
| M: Mechanism explanation | Identifies how a workflow process produces the failure, beyond restating the symptom. | “Too few lemmas; increase ” merely redescribes a shortage. |
| D: Diagnostic check | Actually checks whether the proposed mechanism occurs or distinguishes competing causes; the target is not only aggregate task performance. | Comparing solve rates at and alone tests settings, not the cause of failure. |
| L: Evidence-guided intervention | The intervention is explicitly justified by diagnostic findings; contradictory or inconclusive evidence can change or halt the proposal. | Increasing regardless of the diagnostic result. |
| Label | Decision rule | Binary group |
|---|---|---|
| F: Direct repair | Observable behavior is limited to immediate repair or settings search, without an upstream mechanism explanation. | First-order |
| H: Mechanism hypothesis | M is present, but no diagnostic check D is carried out. | Second-order |
| I: Mechanism investigation | M is present and D is carried out; findings may be inconclusive or refute the hypothesis. | Second-order |
| G: Evidence-guided intervention | M and D are present, and the resulting intervention satisfies L. Task-level success is not required for this label. | Second-order |
| U: Insufficient information | The available record does not support a judgment of the attempt’s inquiry process. | Second-order, by convention |
| Observed optimization behavior | Label | Rationale |
|---|---|---|
| Claims that more lemmas are needed; raises from 10 to 20 and observes a higher solve rate. | F | Establishes a setting’s performance, not why the failure occurred. |
| Later claims that there are too many lemmas; lowers from 20 to 15 and observes a higher solve rate. | F | Opposite settings search; evaluation across instances does not itself establish mechanism inquiry. |
| Hypothesizes that lexical ranking pushes a key lemma below top- , then replaces the retriever without checking that explanation. | H | Provides a process-level hypothesis but no diagnostic check. |
| Checks the needed lemma’s rank in failed cases and whether it actually entered the generation prompt. | I | Tests whether key evidence was omitted by the retrieval process. |
| Holds the count fixed while replacing an irrelevant lemma with a needed one; or preserves needed lemmas while removing irrelevant ones. | I | Probes evidence quality or distraction rather than only searching over counts. |
| Uses diagnostic findings to select ranking, filtering, or count changes, then evaluates the intervention. | G | The intervention follows from mechanism evidence rather than a predetermined solution. |
| Label | Attempts | Share |
|---|---|---|
| F: Direct repair | 86 | 64.2% |
| H: Mechanism hypothesis | 29 | 21.6% |
| I: Mechanism investigation | 8 | 6.0% |
| G: Evidence-guided intervention | 7 | 5.2% |
| U: Insufficient information | 4 | 3.0% |
| Total | 134 | 100% |
| Trajectory | Initial | Final | Cost | |
|---|---|---|---|---|
| 1 | 29.4 | 46.2 | +16.8 | 25 |
| 2 | 29.4 | 50.4 | +21.0 | 28 |
| 3 | 29.4 | 48.0 | +18.6 | 26 |
| 4 | 29.4 | 44.9 | +15.5 | 23 |
| 5 | 29.4 | 51.6 | +22.2 | 32 |
| Mean / total | 29.4 | 48.2 | +18.8 | 134 |
| Setting | Lean proving | Verus VCG |
|---|---|---|
| Benchmark / subset | RSI-Exam + VeriSoftBench | VeriContest, hard |
| Training / test problems | 92 / 48 | 113 / 46 |
| Outer improvement model | GPT-5.6-sol | GPT-5.6-sol |
| Inner model snapshot | gpt-5.4-2026-03-05 | gpt-5.4-2026-03-05 |
| Inner temperature / seed | 0 / 0 | 0 / 0 |
| Outer budget per run | 24 hours | 24 hours |
| Method | Tested | Revised | Accepted G |
|---|---|---|---|
| Lean proving | |||
| Naive RSI | 273/1840 | 16/1840 | 34/62 |
| SO-RSI | 1482/3223 | 729/3223 | 51/76 |
| w/o probes | — | — | 39/69 |
| w/o memory | 1194/2870 | 438/2870 | 43/72 |
| Verus VCG | |||