Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
Organizations: OnCorps
Abstract
An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches without touching a weight, so frozen weights buy an eventual ceiling but no stationarity along the way. The criterion also separates three regimes usually merged: search within a fixed class, test-time training that raises the ceiling itself, and scaffold rewriting between them. Audit the scaffold, not the checkpoint. The same ceiling binds sideways. Best-of- orchestration realizes the best worker's ceiling exactly: width buys rate, not budget. Re-consulting a fixed pool has a horizon computable in advance, decided by the pool alone, and the one arrangement that would beat it, a weighted vote, needs diversity real workers lack: on 30 same-family workers the failure overlap sits at its maximum, and a majority fails 23/55 (42%) of tasks. We obtain the criterion by reading refinement as gradient boosting on the residual error between draft and target, a patch or git diff, and then measuring where that reading breaks: patches compose instead of standing beside each other to be voted on, and failures overlap. What we measure is saturation. Per-round improvement decays toward zero on SWE-bench, and churn decays geometrically across 401 production sessions, a shape shared with a pre-AI human baseline that establishes the regime without identifying its cause. Both breaks are engineering choices rather than laws about code, so together they specify a harness worth building.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Boosting concept | Agentic coding analog | Concrete example |
|---|---|---|
| Input | Specification | Prompts, GitHub tickets |
| Target | Optimal Code | The optimal “best” code for |
| Current State | Current Draft | The code generated by previous agents |
| Residual | Code Edit | A git diff or patch |
| Weak learner | Agent Patch Generator | LLM call proposing an edit |
| Loss Function | Structured Distance | Token edit distance, AST distance |
| Effect | df | partial | Cohen’s | ||
|---|---|---|---|---|---|
| Model (Sonnet vs. Haiku) | 53.46 | (1, 54) | ¡0.0001 | 0.497 | 0.99 (huge) |
| Feedback mode | 0.50 | (2, 108) | 0.61 | 0.009 | 0.10 (small) |
| Model mode | 0.81 | (2, 108) | 0.45 | 0.015 | 0.12 (small) |
| Commit floor | Pairs | AI higher | Sign | Wilcoxon | Median |
|---|---|---|---|---|---|
| none | 21 | 14 | 0.19 | 0.022 | +0.127 |
| 14 | 10 | 0.18 | 0.011 | +0.094 | |
| 9 | 6 | 0.51 | 0.098 | +0.061 | |
| 7 | 5 | 0.45 | 0.11 | +0.127 |
| Result | Lean declarations |
|---|---|
| Thm. 10 , training-error bound | thm1_zero_one_loss_bound (the bound as stated); thm1_training_error_bound for the underlying exponential-loss identity , with z_decomposition , indicator_le_exp , exp_loss_pos |
| Thm. 13 , forward stagewise equivalence | thm4_fsam_optimal_alpha ; thm4_fsam_global_minimum , exp_half_log |
| Cor. 11 , exponential decay | cor2_exponential_decay ; z_le_sqrt_edge , sqrt_edge_le_exp , z_bound_from_edge |
| Thm. 2 , monotone-improvement refinement game | thm5_refinement_game ; thm5_saturation |
| Cor. 15 , high-probability refinement | corollary7_high_prob_refinement ; prob_inter_bound , sum_telescope . Verified against Mathlib’s measure-theoretic library ( IsProbabilityMeasure and the finite union bound) |
| Prop. 3 , stationarity dichotomy | prop8_saturation_forces_decay (part i), prop8_sustained_edge_diverges (part ii) |
| Result | Lean declarations |
|---|---|
| Prop. 5 , best-of- ceiling equality | prop15_tight_ceiling_orchestration (tight equality); prop15_orchestration_ceiling , prop15_orchestration_dominates for the weaker upper-bound-hypothesis version |
| Prop. 17 , weighted aggregation beats the weak-learning baseline | cor_soft_aggregation_beats_weak_baseline ; built on the pure-analysis lemma exists_ensemble_size_beating_weak_threshold |
| Prop. 18 , multi-round best-of- | thm_width_refinement_game ; cor_width_dominates_best_worker , cor_width_saturation , with exists_width_orchestrator for non-vacuity |
| Prop. 6 , bounded overlap characterizes a correct vote | thm_overlap_iff_zero_error (characterization); cor_tight_overlap_iff ( form), thm_bounded_overlap_zero_error (sufficient direction), thm_overlap_boundary_sharp with exists_boundary_overlap_family (sharpness at the boundary), cor_overlap_markov_bound (counting-only baseline) |
| Reweighted edge hypothesis via the density ratio | lem_bounded_reweighting_edge takes the concentration cap as given; D_le_of_margin_bound and cor_rho_from_confidences derive it from logged confidences, and cor_reweighting_edge_from_confidences states the -free conclusion. The sources spell as rho and as B , the names they carried before this paper separated those glyphs |
| Prop. 19 , immediate reuse zeroes the edge | thm_self_reuse_zero_edge ; rests on the derived reweighting recursion D_succ_eq and reused_worker_err_eq , with exists_zero_edge_reuse_instance for non-vacuity |