Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation
Organizations: College of Aeronautics and Engineering, Kent State University
Abstract
Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the chosen programs help. This study examines this distinction using 32 layer-skipping and repetition programs on two models and 4,413 multiple-choice items. The analysis compares their gains over a fixed action selected without the evaluation prompt with those of input-blind perturbations at the same sites, re-evaluating selections on another prompt. With shared option order, the controls give 10.2-11.8 and 15.6-19.4 percentage points of headroom on Qwen3-4B-Base and Llama-3.1-8B, exceeding the real programs' 9.0 and 10.1 in all three random-direction draws per model. They match answer-change rate only, and the ordering depends on the menu: in post hoc comparisons, real programs lead on Llama's repeat-only menu in every draw. A smaller KL-calibrated comparison, including an input-dependent control, favours real programs in point estimate, with inconclusive corrected tests. Fixed letter offsets produce headroom of similar scale. Rotating options sharply reduces both families' headroom, while leaving positive real-minus-control differences of 1.4-2.3 and 3.7-4.5 points; their magnitudes and statistical support depend on further adjustments and the reference. A supplementary generated-answer test finds that search-selected programs keep a 26.0-point advantage over programs selected for other problems after rewording, without a placebo comparison. These results show that substantial headroom can persist across prompts with shared option order without establishing a benefit specific to the selected layer computation; neither ordering against these controls identifies that benefit.
Figures & tables
| Property | RD-fixed | RD-item | SH | Rot |
|---|---|---|---|---|
| Study menu | 32 programs | 3 programs | 3 programs | 3 programs |
| Effect-size match | answer changes | KL | KL | KL |
| Direction shared by all items | yes | no | no | no |
| Same perturbation on shared text | yes | yes | yes | no |
| Model-generated update | no | no | other input’s | own, rotated |
| Reads current input | no | no | no | yes |
| Same prompt | G repro : tie average | : kept | ||||
| Family | Fixed | Rotated | Fixed | Rotated | Fixed | Rotated |
| Qwen3-4B-Base | ||||||
| Real | 29.1\,{\color[rgb]{0.22,0.22,0.22}[28.0,30.3]} | \mathbf{29.3}\,{\color[rgb]{0.22,0.22,0.22}[28.3,30.3]} | 9.0\,{\color[rgb]{0.22,0.22,0.22}[8.1,9.9]} | \mathbf{-3.3}\,{\color[rgb]{0.22,0.22,0.22}[-4.0,-2.7]} | 11.4\,{\color[rgb]{0.22,0.22,0.22}[10.5,12.3]} | \mathbf{1.3}\,{\color[rgb]{0.22,0.22,0.22}[0.7,2.0]} |
| RD-fixed | \mathbf{29.2}\,{\color[rgb]{0.22,0.22,0.22}[28.1,30.4]} | 29.0\,{\color[rgb]{0.22,0.22,0.22}[28.0,30.0]} | \mathbf{11.8}\,{\color[rgb]{0.22,0.22,0.22}[10.9,12.8]} | -5.7\,{\color[rgb]{0.22,0.22,0.22}[-6.3,-5.1]} | \mathbf{13.9}\,{\color[rgb]{0.22,0.22,0.22}[13.0,14.8]} | -0.7\,{\color[rgb]{0.22,0.22,0.22}[-1.3,-0.1]} |
| -0.1\,{\color[rgb]{0.22,0.22,0.22}[-1.0,0.8]} | 0.3\,{\color[rgb]{0.22,0.22,0.22}[-0.5,1.1]} | -2.8\,{\color[rgb]{0.22,0.22,0.22}[-3.8,-1.9]} | 2.3\,{\color[rgb]{0.22,0.22,0.22}[1.6,3.0]} | -2.5\,{\color[rgb]{0.22,0.22,0.22}[-3.5,-1.6]} | 2.0\,{\color[rgb]{0.22,0.22,0.22}[1.3,2.7]} | |
| Llama-3.1-8B | ||||||
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Unmodified pass, action family, and number of edited candidates (excluding ). | |
| Correct-letter margin and 0/1 correctness for item , prompt and action . | |
| Best-accuracy tie set on selection prompt ; static action selected without evaluation outcomes. | |
| Headroom over the static action when selection and scoring use the same prompt, or different prompts. | |
| Fixed-order half-headroom contrast: . | |
| Real minus control headroom; control divided by real headroom. |
| Action | Placebo | Mean difference (nats) placebo real; ref. 0 | Noise variance ratio placebo / real; ref. 1 |
|---|---|---|---|
| skip 16–19 | RD-item | +0.19\,{\color[rgb]{0.22,0.22,0.22}[+0.14,+0.24]} | 0.80\,{\color[rgb]{0.22,0.22,0.22}[0.72,0.91]} |
| SH | +0.16\,{\color[rgb]{0.22,0.22,0.22}[+0.11,+0.21]} | 1.08\,{\color[rgb]{0.22,0.22,0.22}[0.91,1.26]} | |
| Rot | +0.11\,{\color[rgb]{0.22,0.22,0.22}[+0.06,+0.16]} | 0.96\,{\color[rgb]{0.22,0.22,0.22}[0.85,1.08]} | |
| repeat 16–19 | RD-item | -0.11\,{\color[rgb]{0.22,0.22,0.22}[-0.16,-0.06]} | 0.92\,{\color[rgb]{0.22,0.22,0.22}[0.83,1.02]} |
| SH | -0.05\,{\color[rgb]{0.22,0.22,0.22}[-0.09,-0.004]} | 0.71\,{\color[rgb]{0.22,0.22,0.22}[0.65,0.78]} | |
| Rot | -0.08\,{\color[rgb]{0.22,0.22,0.22}[-0.13,-0.04]} | 0.85\,{\color[rgb]{0.22,0.22,0.22}[0.76,0.94]} |
| Action | Placebo | Target KL | Scale | KL ratio (placebo / real) | |
|---|---|---|---|---|---|
| Calibration | Evaluation | ||||
| skip 16–19 | RD-item | ||||
| SH | |||||
| Rot | |||||
| repeat 16–19 | RD-item | ||||
| SH | |||||
| Analysis | Role and specification relative to outcomes |
|---|---|
| Initial half-headroom and equivalence hypotheses | Specified before development outcomes; the original joint power requirement stopped evaluation. No equivalence claim is made. |
| Half-headroom confirmation | Confirmation specified after development estimates and before untouched evaluation; the threshold, menu and scoring rules were retained. |
| Original descriptive checks | Unmodified-pass reference, original letter-shift diagnostic, menu-size curves and cluster sensitivity were planned. |
| Rotation and other diagnostic extensions | Exploratory relative to the original study: rotated options, alternative ties, fitted letter offsets, decoupled and mixture references, covariance tests, and later paired decompositions. |
| Two fresh RD draws | Robustness contrasts specified after original findings and before the new draws; evaluation items were reused. |
| Three-draw means and expanded submenus | Post hoc summaries of stored outputs, with conditional pointwise intervals and no new confirmation rule. |
| (pp): chosen static | (pp): static | ||||
| Task | (%) | Real | RD-fixed | Real | RD-fixed |
| Qwen3-4B-Base | |||||
| MMLU-Pro | |||||
| BBH | |||||
| ARC-Challenge | |||||
| Llama-3.1-8B | |||||
| Menu ( ) | Draw | Real | RD-fixed | ||
|---|---|---|---|---|---|
| Qwen3-4B-Base | |||||
| Repeats (16) | 1 | 7.3 [6.5, 8.2] | 7.2 [6.4, 8.0] | 0.1 [−0.7, 1.0] | 3.5 [2.8, 4.3] |
| 2 | 7.3 [6.5, 8.2] | 8.8 [7.9, 9.6] | −1.4 [−2.3, −0.5] | 5.1 [4.3, 5.9] | |
| 3 | 7.3 [6.5, 8.2] | 9.9 [9.0, 10.8] | −2.6 [−3.5, −1.7] | 6.3 [5.5, 7.0] | |
| Skips (16) | 1 | 5.6 [4.8, 6.5] | 9.8 [8.8, 10.7] | −4.1 [−5.1, −3.2] | 6.9 [6.1, 7.8] |
| 2 | 5.6 [4.8, 6.5] | 7.5 [6.6, 8.3] | −1.8 [−2.7, −1.0] | 4.6 [3.9, 5.4] | |
| Menu ( ) | Draw | Real | RD-fixed | |
|---|---|---|---|---|
| Qwen3-4B-Base | ||||
| Repeats (16) | 1 | −1.1 [−1.7, −0.5] | −2.5 [−3.0, −1.9] | 1.4 [0.7, 2.1] |
| 2 | −1.1 [−1.7, −0.5] | −2.4 [−2.9, −1.8] | 1.3 [0.6, 2.0] | |
| 3 | −1.1 [−1.7, −0.5] | −2.2 [−2.8, −1.7] | 1.1 [0.5, 1.8] | |
| Skips (16) | 1 | −6.7 [−7.3, −6.0] | −8.4 [−9.0, −7.8] | 1.7 [1.0, 2.4] |
| 2 | −6.7 [−7.3, −6.0] | −7.6 [−8.2, −6.9] | 0.9 [0.2, 1.6] | |
| Model | Source family | Source headroom (pp) | Offset headroom (pp) | Offset / source (%) |
|---|---|---|---|---|
| Qwen3-4B-Base | Real | 9.0 [-1pt][8.1, 9.9] | 8.8 [-1pt][8.0, 9.6] | 98.1 [-1pt][87.5, 109.6] |
| RD-fixed | 11.8 [-1pt][10.9, 12.8] | 12.2 [-1pt][11.3, 13.1] | 103.2 [-1pt][95.3, 112.0] | |
| Llama-3.1-8B | Real | 10.1 [-1pt][9.1, 11.0] | 10.9 [-1pt][10.1, 11.8] | 108.6 [-1pt][99.0, 119.4] |
| RD-fixed | 15.6 [-1pt][14.5, 16.6] | 13.8 [-1pt][12.9, 14.8] | 88.8 [-1pt][83.1, 94.8] |
| Fixed: (pp) | Rotated: (pp) | |||||
| Static | Real | RD-fixed | Real | RD-fixed | ||
| Qwen3-4B-Base | ||||||
| Dev. | \mathbf{7.3}\,{\color[rgb]{0.22,0.22,0.22}[6.4,8.2]} | 7.0\,{\color[rgb]{0.22,0.22,0.22}[6.2,7.8]} | 0.3\,{\color[rgb]{0.22,0.22,0.22}[-0.6,1.2]} | \mathbf{-3.1}\,{\color[rgb]{0.22,0.22,0.22}[-3.7,-2.4]} | -4.0\,{\color[rgb]{0.22,0.22,0.22}[-4.6,-3.4]} | 0.9\,{\color[rgb]{0.22,0.22,0.22}[0.2,1.6]} |
| Family | 7.6\,{\color[rgb]{0.22,0.22,0.22}[6.7,8.4]} | \mathbf{7.9}\,{\color[rgb]{0.22,0.22,0.22}[7.0,8.7]} | -0.3\,{\color[rgb]{0.22,0.22,0.22}[-1.2,0.6]} | \mathbf{-3.0}\,{\color[rgb]{0.22,0.22,0.22}[-3.6,-2.3]} | -3.7\,{\color[rgb]{0.22,0.22,0.22}[-4.3,-3.1]} | 0.7\,{\color[rgb]{0.22,0.22,0.22}[0.03,1.4]} |
| 7.0\,{\color[rgb]{0.22,0.22,0.22}[6.2,7.8]} | 7.0\,{\color[rgb]{0.22,0.22,0.22}[6.2,7.8]} | 0.0\,{\color[rgb]{0.22,0.22,0.22}[-0.6,0.6]} | \mathbf{-3.5}\,{\color[rgb]{0.22,0.22,0.22}[-4.2,-2.9]} | -4.3\,{\color[rgb]{0.22,0.22,0.22}[-4.9,-3.7]} | 0.8\,{\color[rgb]{0.22,0.22,0.22}[0.3,1.2]} | |
| Llama-3.1-8B | ||||||
| Headroom (pp) | Drop (pp) | Offset-menu (pp) | |||
| Family | Fixed | Rotated | Fixed rotated | Fixed | Rotated |
| Qwen3-4B-Base | |||||
| Real | 9.0\,{\color[rgb]{0.22,0.22,0.22}[8.1,9.9]} | \mathbf{-3.3}\,{\color[rgb]{0.22,0.22,0.22}[-4.0,-2.7]} | 12.3\,{\color[rgb]{0.22,0.22,0.22}[11.5,13.2]} | ||
| RD-fixed | \mathbf{11.8}\,{\color[rgb]{0.22,0.22,0.22}[10.9,12.8]} | -5.7\,{\color[rgb]{0.22,0.22,0.22}[-6.3,-5.1]} | 17.5\,{\color[rgb]{0.22,0.22,0.22}[16.6,18.4]} | ||
| Llama-3.1-8B | |||||
| Real | 10.1\,{\color[rgb]{0.22,0.22,0.22}[9.1,11.0]} | \mathbf{-2.5}\,{\color[rgb]{0.22,0.22,0.22}[-3.2,-1.8]} | 12.5\,{\color[rgb]{0.22,0.22,0.22}[11.6,13.5]} | ||
| Rule | Family | All correct | All wrong | Several | Unique | Total |
|---|---|---|---|---|---|---|
| Qwen3-4B-Base; shared order | ||||||
| Tie average | Real | 0.0 | 0.1 | 6.6 | 2.4 | 9.0 |
| RD-fixed | 0.0 | 0.1 | 8.6 | 3.2 | 11.8 | |
| 0.0 | 0.0 | −2.0 | −0.8 | −2.8 | ||
| kept | Real | 0.0 | 0.0 | 9.0 | 2.4 | 11.4 |
| RD-fixed | 0.0 | 0.0 | 10.7 | 3.2 | 13.9 | |
| Headroom (pp) | Placebo / Real | (pp) | |||
|---|---|---|---|---|---|
| Family | Real placebo | ||||
| Real | — | — | — | ||
| RD-item | 0.95\,{\color[rgb]{0.22,0.22,0.22}[0.87,1.04]} | 0.69\,{\color[rgb]{0.22,0.22,0.22}[0.50,0.94]} | 2.1\,{\color[rgb]{0.22,0.22,0.22}[0.4,3.7]} | ||
| SH | 0.99\,{\color[rgb]{0.22,0.22,0.22}[0.91,1.08]} | 0.79\,{\color[rgb]{0.22,0.22,0.22}[0.61,1.00]} | 1.4\,{\color[rgb]{0.22,0.22,0.22}[0.01,2.9]} | ||
| Rot | 1.03\,{\color[rgb]{0.22,0.22,0.22}[0.95,1.13]} | 0.81\,{\color[rgb]{0.22,0.22,0.22}[0.61,1.11]} | 1.3\,{\color[rgb]{0.22,0.22,0.22}[-0.7,2.9]} | ||
| Accuracy (%) | (pp) | ||||
|---|---|---|---|---|---|
| Rewordings | Unmodified | Own program | Donors | Own donors | |
| All kept rewordings | 26.0\,{\color[rgb]{0.22,0.22,0.22}[20.9,31.4]} | ||||
| Same executed length | 26.8\,{\color[rgb]{0.22,0.22,0.22}[21.3,32.1]} | ||||
| Sentence by sentence | |||||
| Restructured | |||||
| Draw | Calibration | Matching fixed | Matching rotated | cluster lower bound | ||
| Qwen3-4B-Base | ||||||
| 1 | 32/32 | Pass (32/32) | Pass (32/32) | 11.8\,{\color[rgb]{0.22,0.22,0.22}[10.9,12.8]} | 7.3\,{\color[rgb]{0.22,0.22,0.22}[6.5,8.2]} | 5.8 |
| 2 | 32/32 | Pass (32/32) | Pass (32/32) | 10.2\,{\color[rgb]{0.22,0.22,0.22}[9.3,11.1]} | 5.7\,{\color[rgb]{0.22,0.22,0.22}[4.9,6.5]} | 4.4 |
| 3 | 32/32 | Pass (32/32) | Pass (32/32) | 11.4\,{\color[rgb]{0.22,0.22,0.22}[10.4,12.3]} | 6.9\,{\color[rgb]{0.22,0.22,0.22}[6.1,7.7]} | 5.5 |
| Llama-3.1-8B | ||||||
| 1 | 32/32 | Pass (32/32) | Pass (32/32) | 15.6\,{\color[rgb]{0.22,0.22,0.22}[14.5,16.6]} | 10.5\,{\color[rgb]{0.22,0.22,0.22}[9.6,11.5]} | 8.7 |
| Draw | fixed | rotated | Rotated cluster interval | Same-prompt |
| Qwen3-4B-Base | ||||
| 1 | -2.8\,{\color[rgb]{0.22,0.22,0.22}[-3.8,-1.9]} | 2.3\,{\color[rgb]{0.22,0.22,0.22}[1.6,3.0]} | [1.5, 3.1] | 29.2\,{\color[rgb]{0.22,0.22,0.22}[28.1,30.4]} |
| 2 | -1.2\,{\color[rgb]{0.22,0.22,0.22}[-2.1,-0.3]} | 1.4\,{\color[rgb]{0.22,0.22,0.22}[0.7,2.2]} | [0.4, 2.4] | 28.0\,{\color[rgb]{0.22,0.22,0.22}[26.9,29.2]} |
| 3 | -2.4\,{\color[rgb]{0.22,0.22,0.22}[-3.3,-1.5]} | 1.4\,{\color[rgb]{0.22,0.22,0.22}[0.7,2.1]} | [0.5, 2.2] | 28.9\,{\color[rgb]{0.22,0.22,0.22}[27.8,30.1]} |
| Llama-3.1-8B | ||||
| 1 | -5.5\,{\color[rgb]{0.22,0.22,0.22}[-6.6,-4.4]} | 3.7\,{\color[rgb]{0.22,0.22,0.22}[2.8,4.5]} | [2.8, 4.5] | 35.5\,{\color[rgb]{0.22,0.22,0.22}[34.3,36.7]} |
| Model | Contrast | Mean [item interval] | Draw SD | Draw range |
|---|---|---|---|---|
| Qwen3-4B-Base | , shared | 6.6 [6.0, 7.2] | 0.85 | 5.7 to 7.3 |
| , shared | −2.1 [−2.9, −1.4] | 0.85 | −2.8 to −1.2 | |
| , rotated | 1.7 [1.1, 2.3] | 0.55 | 1.4 to 2.3 | |
| Llama-3.1-8B | , shared | 13.0 [12.1, 13.8] | 2.11 | 10.5 to 14.4 |
| , shared | −7.9 [−8.9, −6.9] | 2.11 | −9.3 to −5.5 | |
| , rotated | 4.0 [3.3, 4.8] | 0.41 | 3.7 to 4.5 |