Verification and Self-Improvement in Agentic AI: Foundations and Limits
Abstract
Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs. A performance score does not distinguish these mechanisms. We compare these changes through bounded verification with hidden terminal randomness. A stage specifies admissible transcripts, polynomial bounds, an alternating verification protocol, and a terminal checker. Its native reach uses default support; its closure frontier permits all support already admitted by the interface. Under a uniform pointwise probability gap and task-relative soundness, these are well-defined languages. We prove that independent majority amplification preserves both languages, whereas existential acceptance over random tapes can admit incorrect outputs. Exact verification is the zero-randomness case, with placement and completeness results. The randomized-verifier classes satisfy ; strict enlargement and depth separation require explicit complexity assumptions, while yields exact companions with the same frontiers. Representation analysis separates invariant acceptance from core-versus-support labels that can change under refactoring. For recursive self-improvement, uniformly bounded self-modification under a common sound interpreter and fixed verification protocol remains within the same verification class. A separate conditional-error budget controls false selection across adaptively chosen candidates. A quota-enforced XOR-synthesis family separates unbounded ratios of search success from changes in the accepted languages; exact and probabilistic audits check the resulting evidence requirements. The framework ties self-improvement claims to obligations on correctness, admissible evidence, verification resources, and selection error.
Figures & tables
| Terminal rule | Incomplete patch | Full repair | Sound for |
|---|---|---|---|
| Four declared tests | Accept | Accept | No |
| All eight declared inputs | Reject | Accept | Yes |
| Task-bound complete certificate | Reject | Accept | Yes |
| Permitted edit | Effect in the XOR family | Contract to retain or establish |
|---|---|---|
| Full-support proposal policy | Changes discovery probability; unchanged | Same exposure, output binding, and exact core. |
| Default quota | Sets equal to the existing | Both defaults belong to the same admitted quota range. |
| Maximum quota | gains weight- targets | New exposure bound and an all-source-witness exclusion proof; core remains sound. |
| Replace the checker | Depends on its accepted behavior | Re-establish gap and task soundness; syntax admission is insufficient. |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Research line | Reported mechanism | Stage comparison and evidence needed |
|---|---|---|
| Autonomy and inherited improvement | Duan et al. organize RSI by control over execution, strategy, experience, deployment adaptation, and inherited improvement mechanisms. They separate structural from effective recursion [ 1 ] . | An autonomy level does not determine . Record accepted-language changes separately from proposal efficiency and successor quality. A fixed enclosing class can accommodate both inherited meta-updates and empirical gains. |
| Prompt and context adaptation | GEPA reflects on execution traces to evolve prompts; ACE incrementally maintains playbooks; RSIAgent carries frozen memory into evaluation [ 21 , 23 , 14 ] . | Declare which context was already available through optional support. Compare memory admissibility and use under matched task and output contracts. |
| Workflow search | AFlow uses Monte Carlo tree search over code-represented workflows and execution feedback [ 22 ] . | Separate a better policy over existing traces from a changed trace grammar. Adding a critic node does not establish universal challenge semantics. |
| Agent and harness code | DGM retains an archive of self-modified agents. AHE exposes editable components and records predicted effects; Self-Harness validates proposed edits with regression tests [ 12 , 20 , 26 ] . | Version the permitted programs, tool behavior, and terminal checker. An edit may affect exposure, acceptance, or only search efficiency; its file location does not determine its semantic role. |
| Editable improvers | STOP optimizes the improver’s own program. Hyperagents jointly exposes task and meta-agent code to modification. AIDE 2 explores research-agent improvement across successive versions [ 18 , 19 , 13 ] . | Version , , and separately. Measure the updated improver on common starting agents and budgets, and identify any fixed outer search or evaluation process. |
| Search evaluation | SIFT aggregates pairwise judge preferences through a Bradley–Terry model to guide parent selection and evaluation priority [ 24 ] . | Treat ranking feedback as a search signal unless its separate acceptance contract is established. Include all selection calls in the error and cost accounting. |