Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what extent a model should adapt as a joint prediction and decision problem. Rather than searching over candidate programs anew for each learning episode, a compiler learns from prior adaptations to predict a vector-valued counterfactual response surface over candidate programs---their expected effects on acquisition, transfer, boundedness, and preservation---and selects a program before adaptation begins. Because this predicted geometry captures multiple behavioral consequences rather than a single winner or scalar score, it can be reused under different downstream priorities without retraining. Across five learning types, preferred programs vary meaningfully across episodes, and this variation is predictable from pre-adaptation information. On Llama-3.1-8B, compiler-selected programs approach exhaustive search while outperforming global and objective-specific defaults. Replication on Gemma-2-9B preserves program heterogeneity and selection headroom, but shows that exploiting this headroom requires accounting for uncertainty when departing from strong defaults. Together, these results show that adaptation search can be amortized across related learning problems, turning prior adaptation experience into a basis for deciding how future learning should occur.
Figures & tables
Figure 1: Selection headroom in adaptation geometry. (A) Fraction of held-out episodes for which each adaptation program is oracle-optimal, shown by learning objective. (B) Episode-specific headroom beyond the objective-fixed policy, measured as oracle utility minus objective-fixed utility. Gray points denote episodes; black points and error bars show the mean and 95% CI. Headroom is largest for lexical binding and factual association and nearly absent for procedural reasoning.
Figure 2: Predicting adaptation geometry from pre-adaptation episode–model signals. (A) Predicted and observed balanced utility across held-out episode–program pairs. (B) Mean absolute error for acquisition, transfer, boundedness, and preservation under the primary predictor and non-episode-conditioned baselines. (C) Within-episode program-ranking fidelity measured by Spearman correlation, pairwise ranking accuracy, and tie-aware top-1 and top-2 oracle recovery.
Figure 3: Program selection from predicted adaptation geometry. (A) Realized balanced utility for global-fixed, objective-fixed, compiler, and oracle selection; values below the axis show mean oracle regret. (B) Oracle regret under alternative utility specifications, using the same predicted geometry without retraining. Percentages show tie-aware top-1 oracle recovery.
Figure 4: Scope and boundary conditions of adaptation compilation. (A) Episode representations alone support selection comparable to the full representation, while module probes and frozen-behavior features are weaker in isolation; the outlined bar marks the primary full representation. (B) Compilation reduces oracle regret when learning families are represented during meta-training, but this advantage reverses when an entire family is withheld (LOFO). (C) Compilation improves over objective-fixed defaults on Llama, whereas on Gemma direct selection from predicted geometry does not improve over the strong objective-fixed policy despite remaining episode-level headroom. Error bars show 95% episode-bootstrap CIs where applicable. Lower oracle regret is better.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Gradient Accum.
Acquisition
Transfer
Boundedness
Preservation
8
0.408
0.240
0.393
0.997
4
0.648
0.587
0.500
0.992
2
0.896
0.867
0.547
0.971
1
0.904
0.853
0.700
0.876
Appendix
Table 1: Optimization calibration on meta-training episodes. Values are averaged across 25 episodes.
Objective
Objective-fixed
Global U
Obj. U
Oracle U
Global Regret
Obj. Regret
Global Opt.
Obj. Opt.
Overall
—
.569
.600
.618
.049
.018
.41
.64
Behavioral
Late
.714
.747
.753
.039
.007
.15
.45
Causal
Middle
.489
.489
.507
.018
.018
.75
.75
Factual
Late
.504
.564
.591
.087
.027
.10
.40
Lexical
Early
.688
.751
.790
.102
.039
.10
.65
Procedural
Middle
.449
.449
.450
.001
.001
.95
.95
Appendix
Table 2: Selection headroom by learning objective. Utilities are measured on held-out test episodes. Global- and objective-fixed programs are selected using meta-training data only. Regret is measured relative to the per-episode oracle, and “Opt.” denotes the fraction of episodes on which the corresponding fixed policy is oracle-optimal.
Objective
Early
Middle
Late
Full-r4
Tie Rate
Mean Margin
Median Margin
Overall
.165
.410
.210
.215
.01
.054
.041
Behavioral
.000
.150
.450
.400
.00
.016
.007
Causal
.000
.750
.150
.100
.00
.072
.059
Factual
.175
.100
.400
.325
.05
.045
.023
Lexical
.650
.100
.000
.250
.00
.077
.076
Procedural
.000
.950
.050
.000
.00
.059
.053
Appendix
Table 3: Oracle-program distribution and winner separation. Program columns report the fraction of held-out episodes for which each configuration is oracle-optimal, using fractional credit for ties. The final columns report the tie rate and the mean and median utility margin between the highest- and second-highest-utility configurations.
Figure 5: Separation between oracle-optimal and second-best programs. Points show the utility margin between the best and second-best adaptation programs for individual held-out episodes, grouped by learning objective. Summary markers indicate the objective-level mean.
Figure 6: Adaptation stochasticity across programs. (A) Variation in utility across three adaptation seeds for each episode–program pair. (B) Stability of the oracle-optimal program across seeds, using both strict unique-winner and tie-aware criteria.
Figure 7: Configuration sensitivity across adaptation outcomes. Acquisition, transfer, boundedness, and preservation under each budget-matched adaptation program, grouped by learning objective. Values summarize held-out test episodes using seed-averaged outcomes.
Objective
Geometry MAE
Acquisition
Transfer
Boundedness
Preservation
Spearman
Pairwise
Behavioral
.015
.023
.030
.007
.002
.900
.917
Causal
.033
.050
.051
.030
.002
.940
.958
Factual
.041
.045
.053
.061
.003
.603
.758
Lexical
.086
.090
.138
.112
.004
.630
.783
Procedural
.025
.022
.001
.073
.002
.940
.958
Appendix
Table 4: Geometry prediction by learning objective. Geometry MAE is averaged across acquisition, transfer, boundedness, and preservation. Spearman correlation and pairwise accuracy measure agreement between predicted and observed program rankings within each episode. The same primary predictor is evaluated across all objectives without access to objective identity.
Figure 8: Geometry prediction across learning objectives. (A) Episode-level mean absolute error between predicted and observed adaptation geometry, grouped by learning objective. Gray points denote held-out episodes and black markers indicate means with 95% bootstrap confidence intervals. (B) Within-episode agreement between predicted and observed program orderings, measured by Spearman rank correlation and pairwise ranking accuracy. Points denote individual episodes and summary markers indicate means with 95% bootstrap confidence intervals. The same predictor is evaluated across all objectives without access to objective identity.
Realized Utility
Oracle Regret
Top-1
Objective
Global
Obj.
Compiler
Global
Obj.
Compiler
Global
Obj.
Compiler
Behavioral
.714
.747
.753
.039
.007
.001
.15
.45
.70
Causal
.489
.489
.504
.018
.018
.004
.75
.75
.90
Factual
.504
.564
.585
.087
.027
.006
.10
.40
.65
Lexical
.688
.751
.746
.102
.039
.044
.10
.65
.65
Procedural
.449
.449
.448
.001
.001
.002
.95
.95
.95
Appendix
Table 5: Program selection by learning objective. Realized utility and oracle regret are reported for the global-fixed, objective-fixed, and compiler policies. Top-1 denotes the fraction of held-out episodes for which each policy selects an oracle-optimal program.
Figure 9: Program-selection performance by learning objective. (A) Oracle regret under global-fixed, objective-fixed, and compiler selection. (B) Fraction of held-out episodes for which each policy selects an oracle-optimal program. Compiler performance is strongest for behavioral, causal, factual, and procedural episodes, while lexical binding remains the principal residual failure case.
Utility
Compiler U
Oracle U
Compiler Regret
Obj.-fixed Regret
Regret Reduction
Switch
Balanced
.607
.618
.011
.018
38.9%
—
Transfer-heavy
.567
.578
.011
.025
53.4%
14%
Boundedness-heavy
.602
.606
.004
.025
84.2%
13%
Preservation-heavy
.683
.692
.009
.015
39.1%
1%
Appendix
Table 6: Recompilation under alternative utility specifications. The same predicted adaptation geometry is evaluated under each utility without retraining the predictor. Regret reduction is measured relative to the objective-fixed policy. “Switch” reports the fraction of compiler-selected programs that differ from the balanced-utility selection.
Representation
Val. MAE
Test MAE
Utility Corr.
Spearman
Pairwise
Top-1
Top-2
Oracle Regret
Episode
.0434
.0399
.9633
.8478
.8983
.83
.96
.0088
Full
.0435
.0400
.9620
.8027
.8750
.77
.93
.0113
Module probes
.0526
.0530
.9348
.7655
.8600
.74
.89
.0141
Frozen behavior
.0563
.0538
.9176
.7406
.8400
.71
.91
.0186
Appendix
Table 7: Complete representation-ablation results. Each restricted predictor is selected using validation performance and evaluated on the same held-out test episodes. Geometry MAE measures prediction error across acquisition, transfer, boundedness, and preservation. Ranking metrics are computed across candidate programs within each episode.
Figure 10: Representation ablations. We compare the full episode–model representation with three restricted feature sets. (A) Mean absolute error in predicted adaptation geometry. (B) Oracle regret of the program selected from predicted geometry. (C) Tie-aware top-1 recovery of an oracle-optimal program. The episode representation alone retains performance comparable to the full representation, while module-level probes and frozen-model behavioral statistics are weaker when used independently. The outlined marker denotes the full representation used as the primary predictor in Experiments 2 and 3.
Representation
Hidden-state representation
Frozen-model loss
Module diagnostics
Backward pass required
Episode
✓
–
–
No
Frozen behavior
–
✓
–
No
Module probes
–
–
✓
Yes
Full
✓
✓
✓
Yes
Appendix
Table 8: Feature sets used in the representation ablation. Configuration descriptors are provided to all predictors; rows describe the episode–model information available in each condition.
Geometry MAE ↓
Oracle Regret ↓
Objective
Episode
Full
Probe
Frozen
Episode
Full
Probe
Frozen
Behavioral
.0153
.0155
.0155
.0167
.0005
.0006
.0005
.0009
Causal
.0346
.0335
.0450
.0323
.0039
.0039
.0092
.0074
Factual
.0379
.0406
.0679
.0501
.0007
.0059
.0236
.0345
Lexical
.0864
.0860
.1100
.1062
.0370
.0438
.0370
.0423
Procedural
.0252
.0245
.0267
.0639
.0021
.0021
.0000
.0078
Appendix
Table 9: Representation ablations by learning objective. Geometry MAE measures prediction error across acquisition, transfer, boundedness, and preservation; oracle regret measures the downstream quality of the program selected from each predicted geometry. Bold values indicate the best result within each objective and metric group.
Figure 11: Generalization beyond represented learning families. (A) Geometry prediction error for held-out episodes from learning families represented during meta-training compared with leave-one-family-out (LOFO) evaluation, where the test objective is excluded from both training and validation. (B) Change in oracle regret of the LOFO compiler relative to the global-fixed policy; positive values indicate improved selection and negative values worse selection.
Held-out family
Val. MAE
Test MAE
Utility Corr.
Spearman
Pairwise
Top-1
Top-2
Behavioral
.050
.451
.862
.520
.708
.30
.60
Causal
.044
.295
.646
.470
.692
.10
.55
Factual
.043
.409
-.110
-.100
.483
.20
.45
Lexical
.026
.341
-.026
-.550
.258
.05
.15
Procedural
.046
.200
-.037
-.050
.475
.05
.60
Macro
.042
.339
.267
.058
.523
.14
.47
Appendix
Table 10: Leave-one-family-out geometry prediction. For each fold, the indicated learning family is excluded entirely from both meta-training and validation. Validation MAE is measured on represented learning families, while test metrics are computed on episodes from the held-out family. Ranking metrics compare predicted and observed program utilities within each episode.
Figure 12: Program-ranking fidelity under leave-one-family-out generalization. Within-episode agreement between predicted and observed adaptation-program orderings is compared when the learning family is represented during meta-training versus excluded entirely under LOFO evaluation. ( A ) Spearman rank correlation. ( B ) Pairwise ranking accuracy. Ranking fidelity decreases across all five objectives when the learning family is unseen, with particularly large degradation for factual, lexical, and procedural episodes.
Held-out family
Global U
Compiler U
Oracle U
Global Regret
Compiler Regret
ΔR
Behavioral
.7144
.7197
.7534
.0390
.0337
+.0052
Causal
.4252
.4252
.5074
.0822
.0822
.0000
Factual
.5044
.5244
.5909
.0865
.0665
+.0201
Lexical
.6882
.6332
.7898
.1016
.1566
-.0550
Procedural
.4493
.3635
.4501
.0008
.0866
-.0858
Macro
.5563
.5332
.6183
.0620
.0851
-.0231
Appendix
Table 11: Program selection under leave-one-family-out generalization. Utilities are realized under balanced utility. Regret is measured relative to the exhaustive episode oracle. ΔR denotes regret reduction relative to the global-fixed policy, such that positive values indicate beneficial zero-shot compilation.
Grad. accum.
Acquisition
Transfer
Boundedness
Preservation
Utility
1
.560
.527
.460
.998
.636
2
.316
.360
.400
.998
.519
4
.168
.207
.220
.998
.398
8
.068
.180
.213
.998
.365
Appendix
Table 12: Gemma optimization-schedule calibration. Values are averaged across the 25 calibration episodes.
Objective
Program
A
T
B
P
U
Behavioral
Early
.087
.494
.000
.999
.395
Middle
.343
.556
.000
.999
.474
Late
.092
.492
.006
.971
.390
Full
.378
.639
.014
.994
.506
Causal
Early
.015
.036
.397
.996
.361
Middle
.057
.072
.422
.997
.387
Appendix
Table 13: Gemma adaptation geometry on held-out test episodes. Values are means over 20 episodes per objective after averaging the three adaptation seeds. A , T , B , and P denote acquisition, transfer, boundedness, and preservation; U is their balanced mean.
Objective
Obj.-fixed
Global U
Obj. U
Oracle U
Global regret
Obj. regret
Global Opt.
Obj. Opt.
Overall
—
.436
.438
.461
.025
.023
.40
.53
Behavioral
Full
.506
.506
.522
.016
.016
.75
.75
Causal
Middle
.388
.387
.399
.011
.012
.30
.55
Factual
Full
.454
.454
.492
.038
.038
.35
.35
Lexical
Late
.524
.536
.572
.048
.036
.10
.50
Procedural
Full
.309
.309
.322
.013
.013
.50
.50
Appendix
Table 14: Gemma selection headroom on the 100 held-out test episodes. Global- and objective-fixed programs are selected using meta-training data only. “Opt.” is the fraction of episodes on which the fixed policy belongs to the oracle set.
Objective
Early
Middle
Late
Full
Overall
.06
.31
.23
.40
Behavioral
.00
.25
.00
.75
Causal
.15
.50
.05
.30
Factual
.00
.20
.45
.35
Lexical
.10
.30
.50
.10
Procedural
.05
.30
.15
.50
Appendix
Table 15: Distribution of oracle-optimal Gemma programs on held-out episodes. Program shares use fractional credit for ties.
Objective
Obj.-fixed
Compiler selections
Compiler U
Obj. U
Oracle U
Compiler regret
Top-1
Top-2
Behavioral
Full
Full: 20
.506
.506
.522
.016
.75
1.00
Causal
Middle
Full: 20
.388
.387
.399
.011
.30
.55
Factual
Full
Full: 20
.454
.454
.492
.038
.35
.85
Lexical
Late
Full: 10; Late: 10
.532
.536
.572
.040
.35
.55
Procedural
Full
Full: 8; Middle: 12
.297
.309
.322
.025
.35
.70
Overall
—
—
.435
.438
.461
.026
.42
.73
Appendix
Table 16: Gemma compiler selection by learning objective. “Selections” reports the number of the 20 held-out episodes assigned to each program. Top-1 and Top-2 are oracle recovery rates.
Statistic
Value
Number of episodes
7
Mean predicted margin
.0027
Median predicted margin
.0031
Mean realized regret
.0501
Appendix
Table 17: Diagnostic for procedural episodes on which the compiler selects middle adaptation while full-stack is oracle-optimal. The predicted margin is the predicted utility advantage that triggers the middle-program selection.
Objective
Llama
Gemma
Behavioral policy
Late
Full
Causal mapping
Middle
Middle
Factual association
Late
Full
Lexical binding
Early
Late
Procedural reasoning
Middle
Full
Appendix
Table 18: Objective-conditioned adaptation defaults across backbones. Defaults are estimated independently from each backbone’s meta-training episodes.