A Generalisation Signal Need Not Be a Model-Selection Signal
Authors: Aditya Nagarsekar, M P Ashish Bhat, Aadi Nesarkar, Vrishti Godhwani, Rahul Yedida, Aditya Challa, Danda Sravan, Snehanshu Saha
Organizations: Department of CS&IS, BITS Pilani, K K Birla Goa Campus · LexisNexis Legal & Professional · Center for AI and Supercomputing, Mahindra University
Model selection in computational biology often relies on validation data drawn from the training regime, even when deployment lies outside it. When validation no longer preserves which model is best, a natural alternative is to rank candidates using properties of the trained network itself. We test this idea using a novel, forward-only proxy motivated by the norm of the Hessian, alongside common Hessian measures, across molecular property, protein fitness, and drug-response tasks. Contrary to our hypothesis, geometry does not become more useful as validation Spearman correlation deteriorates: augmenting validation helps some shifts but significantly harms others. More surprisingly, the proxy still correlates with generalisation gap on most tasks even when Hessian trace and top-eigenvalue relationships are weak or reversed, yet this signal does not reliably identify the deployment-best model. A curvature bound need not preserve cross-model rankings, and low geometric scores can even favour collapsed predictors. Thus, a generalisation signal need not be a model-selection signal.
Figures & tables
Dataset
Metric
Val loss
Proxy
λmax
Random
Caco2-Wang (random)
MAE ↓
0.351 ± 0.013
0.343 ± 0.015
0.363 ± 0.022
0.358 ± 0.016
Caco2-Wang (scaffold)
MAE ↓
0.432 ± 0.047
0.421 ± 0.040
0.442 ± 0.038
0.443 ± 0.039
Lipophilicity (random)
MAE ↓
0.568 ± 0.013
0.571 ± 0.017
0.585 ± 0.021
0.590 ± 0.012
Lipophilicity (scaffold)
MAE ↓
0.656 ± 0.025
0.654 ± 0.028
0.676 ± 0.027
0.681 ± 0.023
Lipophilicity (mismatch)
MAE ↓
0.650 ± 0.022
0.649 ± 0.030
0.663 ± 0.028
0.673 ± 0.025
FLIP2 Amylase
Spearman ρ↑
0.011 ± 0.103
-0.010 ± 0.115
-0.062 ± 0.104
-0.015 ± 0.028
Table 1: Shared-pool deployment performance using each benchmark’s native metric, after the post-hoc degeneracy audit. Proxy is the layer-averaged proxy. Mean ± SD over outer replications is reported. Random is expected uniform selection. NDCG uses the full test set with no cutoff. Per-condition results under all metrics are in Appendix E .
Random
TPE
HEBO
Dataset
search
Val
Val+ Proxy
Val
Val+ Proxy
Caco2-Wang (random)
0.774 ± 0.024
0.774 ± 0.024
0.779 ± 0.026
0.775 ± 0.025
0.777 ± 0.025
Caco2-Wang (scaffold)
0.688 ± 0.085
0.684 ± 0.099
0.684 ± 0.102
0.678 ± 0.089
0.698 ± 0.080
Lipophilicity (random)
0.750 ± 0.013
0.749 ± 0.016
0.754 ± 0.016
0.746 ± 0.024
0.752 ± 0.013
Lipophilicity (scaffold)
0.692 ± 0.031
0.688 ± 0.036
0.705 ± 0.027
0.685 ± 0.021
0.701 ± 0.025
Lipophilicity (mismatch)
0.704 ± 0.027
0.711 ± 0.028
0.713 ± 0.031
0.707 ± 0.028
0.711 ± 0.031
Table 2: Sequential HPO deployment Spearman. Validation+Proxy denotes matched multi-objective search on the pre-registered layer-averaged proxy. Spearman is used throughout for a common scale-free comparison. Mean ± SD over outer replications; random search is the reference. Full per-arm results are in Appendix F .
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
name
deploys the candidate with the lowest
code
Val
validation MSE
val
Proxy
layer-averaged activation proxy, Eq. ( 6 )
fg_legacy
Proxy (pen.)
penultimate-layer proxy, Eq. ( 5 )
fg_penult
λmax
top Hessian eigenvalue of the training MSE
hess_top
Val+X
rank sum of validation MSE and signal X
val+legacy , val+penult , val+hess
Train
training MSE
train_mse
Appendix
Table 3: Selector names, the model each deploys, and the identifier in the code release.
Figure 1: Every selector against blind selection, per condition. Bars below the line indicate a selector that deploys a better model than picking uniformly at random from the same pool. A selector that cannot clear this line carries no usable model-selection information, whatever its correlation with the generalisation gap. Audited pools; Cond7 is Lipophilicity (mismatch) and Hydro is Hydrophobic Core.
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
0.3354 [0.2945, 0.3642]
-0.0292
20/0
< 0.001
< 0.001
Val
0.3575 [0.3283, 0.3936]
–
–
–
–
Val+Proxy
0.3446 [0.3144, 0.3749]
-0.0085
14/2
0.013
0.105
Val+Proxy (pen.)
0.3446 [0.3144, 0.3767]
-0.0105
13/2
0.009
0.081
Val+ λmax
0.3673 [0.3203, 0.3882]
-0.0113
12/5
0.177
0.355
Proxy
0.3490 [0.3178, 0.3723]
-0.0103
15/5
0.033
0.164
Appendix
Table 4: Shared pool, Caco2 (random), audited pools (20 outer runs, median 48/48 candidates kept).
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
0.4448 [0.4168, 0.4744]
-0.0376
18/0
< 0.001
0.002
Val
0.4750 [0.4462, 0.5269]
–
–
–
–
Val+Proxy
0.4766 [0.4347, 0.5090]
+0.0000
7/3
0.047
0.375
Val+Proxy (pen.)
0.4766 [0.4347, 0.5179]
+0.0000
7/5
0.117
0.632
Val+ λmax
0.4782 [0.4549, 0.5277]
+0.0000
7/8
0.496
0.709
Proxy
0.4774 [0.4484, 0.5161]
-0.0189
11/7
0.157
0.632
Appendix
Table 5: Shared pool, Caco2 (scaffold), audited pools (20 outer runs, median 48/48 candidates kept).
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
0.3964 [0.3847, 0.4142]
-0.0110
23/0
< 0.001
< 0.001
Val
0.4089 [0.3963, 0.4259]
–
–
–
–
Val+Proxy
0.4029 [0.3867, 0.4239]
+0.0000
11/2
0.019
0.074
Val+Proxy (pen.)
0.4044 [0.3867, 0.4241]
+0.0000
11/3
0.019
0.074
Val+ λmax
0.4188 [0.4009, 0.4381]
+0.0084
6/20
0.001
0.005
Proxy
0.4113 [0.3927, 0.4313]
-0.0004
15/9
0.775
0.775
Appendix
Table 6: Shared pool, Lipo (random), audited pools (30 outer runs, median 47/48 candidates kept).
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
0.4942 [0.4573, 0.5307]
-0.0114
20/0
< 0.001
0.001
Val
0.5129 [0.4716, 0.5505]
–
–
–
–
Val+Proxy
0.5001 [0.4743, 0.5491]
+0.0000
12/5
0.093
0.371
Val+Proxy (pen.)
0.5001 [0.4759, 0.5505]
+0.0000
12/6
0.102
0.371
Val+ λmax
0.5255 [0.4825, 0.5535]
+0.0123
5/20
0.003
0.013
Proxy
0.5113 [0.4654, 0.5373]
+0.0000
14/10
0.331
0.663
Appendix
Table 7: Shared pool, Lipo (scaffold), audited pools (30 outer runs, median 47/48 candidates kept).
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
0.4768 [0.4621, 0.5161]
-0.0089
19/0
< 0.001
0.001
Val
0.4947 [0.4770, 0.5319]
–
–
–
–
Val+Proxy
0.4775 [0.4721, 0.5304]
-0.0004
10/1
0.013
0.043
Val+Proxy (pen.)
0.4786 [0.4721, 0.5304]
-0.0014
11/2
0.011
0.043
Val+ λmax
0.5159 [0.4853, 0.5509]
+0.0169
4/15
0.003
0.015
Proxy
0.5058 [0.4733, 0.5314]
-0.0016
13/4
0.309
0.618
Appendix
Table 8: Shared pool, Lipo (mismatch), audited pools (20 outer runs, median 47/48 candidates kept).
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
1.4681 [1.3505, 1.6811]
-0.8348
19/0
< 0.001
0.001
Val
2.5773 [2.1102, 2.6980]
–
–
–
–
Val+Proxy
2.3381 [1.8289, 2.5956]
+0.0000
8/2
0.093
0.556
Val+Proxy (pen.)
2.3149 [2.0077, 2.5956]
+0.0000
8/3
0.131
0.653
Val+ λmax
2.6598 [2.5062, 2.8849]
+0.0000
3/6
0.214
0.737
Proxy
2.1202 [1.7053, 2.4845]
-0.3628
14/5
0.027
0.188
Appendix
Table 9: Shared pool, Amylase, audited pools (20 outer runs, median 11/48 candidates kept).
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
20.5821 [20.2789, 20.9410]
-2.2398
30/0
< 0.001
< 0.001
Val
22.9161 [22.3655, 23.3435]
–
–
–
–
Val+Proxy
23.4878 [23.0254, 23.8202]
+0.5885
5/20
< 0.001
0.004
Val+Proxy (pen.)
23.2913 [23.0405, 23.7761]
+0.6543
6/20
0.003
0.020
Val+ λmax
23.4338 [22.4991, 24.2371]
+0.2414
11/19
0.114
0.343
Proxy
22.8071 [22.4965, 23.0753]
-0.0748
16/14
0.792
0.792
Appendix
Table 10: Shared pool, Hydrophobic Core, audited pools (30 outer runs, median 34/48 candidates kept).
selector
median [IQR]
Δ vs Val
w/l
p
pHolm
Oracle
0.6666 [0.4961, 0.7516]
-0.0529
19/0
< 0.001
0.001
Val
0.7086 [0.5330, 0.8320]
–
–
–
–
Val+Proxy
0.7116 [0.5428, 0.8201]
-0.0018
10/8
0.983
1.000
Val+Proxy (pen.)
0.7181 [0.5428, 0.8201]
+0.0000
9/9
0.647
1.000
Val+ λmax
0.6925 [0.5297, 0.8018]
+0.0000
9/7
0.255
1.000
Proxy
0.7256 [0.5899, 0.8203]
+0.0286
8/12
0.177
1.000
Appendix
Table 11: Shared pool, GDSC2, audited pools (20 outer runs, median 48/48 candidates kept).
unfiltered (primary)
audited
condition
Δ
95% CI
w/l/t
pHolm
Δ
95% CI
w/l/t
pHolm
Proxy against Val
Lipo (mismatch)
-0.0009
[-0.0068, 0.0000]
12/5/3
0.407
-0.0016
[-0.0073, 0.0000]
13/4/3
0.618
Caco2 (scaffold)
-0.0189
[-0.0301, +0.0093]
11/7/2
0.328
-0.0189
[-0.0301, +0.0093]
11/7/2
0.471
Amylase
-1.3934
[-1.5327, -1.1955]
20/0/0
< 0.001
-0.3628
[-0.6133, -0.0268]
14/5/1
0.108
Hydrophobic Core
+0.1634
[-0.1005, +0.5858]
12/18/0
0.328
-0.0748
[-0.4902, +0.3885]
16/14/0
0.792
Appendix
Table 12: The pre-registered confirmatory analysis, deployment MSE, Holm-corrected across the four-condition family within each selector. Reported on the original unfiltered pools (the pre-registered primary analysis) and again on the audited pools (post-hoc). Δ is the paired median difference against Val, negative favouring the selector; the interval is a 95% percentile bootstrap over replications; w/l/t counts wins, losses and ties. The signed-rank test discards ties (Appendix B ).
condition
under MSE
under MAE
under ρ
Caco2 (random)
–
–
–
Caco2 (scaffold)
–
–
–
Lipo (random)
–
–
–
Lipo (scaffold)
–
–
–
Lipo (mismatch)
Val+Proxy, Val+Proxy (pen.)
–
–
Amylase
–
–
–
Appendix
Table 13: Metric robustness on the audited pools: selectors that beat Val at Holm-adjusted α=0.05 under each metric, with the selection rule unchanged. MSE is train- σ standardised; MAE is what the TDC leaderboards report; ρ is Spearman, what FLIP2 reports.
Figure 2: Deployment loss against search budget. More search helps where validation ranking transfers and does not where it fails. Unfiltered search trajectories; Cond7 is Lipophilicity (mismatch) and Hydro is Hydrophobic Core.
Val
Val+Proxy
Val+Proxy (pen.)
Val+ λmax
condition
Random
TPE
HEBO
TPE
HEBO
TPE
HEBO
TPE
HEBO
Lipo (mismatch)
0.518
0.503
0.514
0.501
0.494
0.494
0.504
0.522
0.534
Hydrophobic Core
22.79
23.29
23.12
23.76
23.69
23.47
23.53
23.74
22.62
Caco2 (random)
0.351
0.355
0.351
0.340
0.341
0.332
0.335
0.344
0.342
Caco2 (scaffold)
0.464
0.481
0.460
0.443
0.457
0.424
0.445
0.464
0.442
Lipo (random)
0.412
0.409
0.424
0.411
0.411
0.415
0.408
0.407
0.416
Appendix
Table 14: Sequential search, all nine arms, deployment MSE (median over replications, unfiltered trajectories). The reference is validation-only TPE; bold marks arms that Holm-beat it within that condition, underline marks arms that are Holm-worse. Holm is applied over the eight contrasts within each condition. Multi-objective arms deploy by the rank-sum rule of Appendix B .
Figure 3: Dead-unit fraction against deployment loss on the unfiltered pools, with constant predictors marked. The two panels are a contrast, not a pattern. On Amylase (left) 24% of candidates collapse to a constant predictor, and because predicting near the training mean is competitive in MSE on that split, they sit among the lowest-loss models, so the proxy selects one in 80% of replications. On Hydrophobic Core (right) many networks are partially dead but none is constant, and the proxy never selects one. Rates for every selector are in Table 16 . Whether minimising a geometric signal rewards collapse is therefore a property of the condition, not of the signal, which is why it has to be audited per condition rather than assumed.
condition
const.
dead > 0.5
median dead
dead@Proxy
Caco2 (random)
0.0%
0.0%
0.003
0.026
Caco2 (scaffold)
0.0%
0.0%
0.002
0.028
Lipo (random)
0.0%
2.8%
0.008
0.098
Lipo (scaffold)
0.0%
2.5%
0.008
0.095
Lipo (mismatch)
0.0%
2.1%
0.008
0.069
Amylase
24.3%
76.5%
0.675
0.998
Appendix
Table 15: Degeneracy by condition, computed before filtering. “const.” is the fraction of candidates whose test-set predictions have std(y^)<10−6 ; “dead > 0.5” the fraction with more than half their ReLU units inactive on every example of the training proxy subset; “dead@Proxy” the median dead-unit fraction of the model Proxy selects.
selector
constant (test)
constant (train)
undefined Spearman
median dead fraction
Val
0%
0%
0/20
0.482
Proxy
80%
100%
14/20
0.998
Proxy (pen.)
85%
100%
14/20
0.973
λmax
100%
100%
16/20
0.984
Val+Proxy
0%
0%
0/20
0.502
Val+Proxy (pen.)
0%
0%
0/20
0.532
Appendix
Table 16: How often each selector deploys a constant predictor on the unfiltered Amylase pools (20 replications). A model counts as constant when the standard deviation of its predictions is below 10−6 , on the test set (as in the audit) or on the training set. Spearman is undefined when the test predictions are exactly constant. No selector deploys a constant predictor on Hydrophobic Core.
filter
Δ
w/l
pHolm
none (primary analysis)
-1.3934
20/0
< 0.001
constant predictors, test-set flag
-1.3938
20/0
< 0.001
constant predictors, training-set flag
-0.8218
19/1
< 0.001
more than half the units dead
-0.3628
14/5
0.108
full audit (both criteria)
-0.3628
14/5
0.108
Appendix
Table 17: Proxy against Val on Amylase under each candidate filter, deployment MSE, Holm-corrected across the four confirmatory conditions. Only the dead-unit criterion makes the difference non-significant.
condition
selector
Δlog10η
Δlog2w
Δ dropout
Δ wd
Caco2 (random)
Val
+0.110
+0.00
-0.029
+5.8e-04
Caco2 (random)
#params
+0.170
+1.50
-0.056
+7.6e-04
Caco2 (scaffold)
Val
+0.222
+0.50
+0.014
-1.3e-06
Caco2 (scaffold)
#params
+0.171
+2.00
+0.112
+2.0e-03
Lipo (random)
Val
+0.000
+0.00
+0.013
+0.0e+00
Lipo (random)
#params
-0.227
+3.00
+0.044
+9.7e-04
Appendix
Table 18: Median paired hyperparameter difference between the selected model and the deployment oracle in the same pool. Positive Δlog10η means the oracle prefers a larger learning rate. Audited pools.
Figure 4: Top Hessian eigenvalue against deployment MSE, on the fixed-architecture subset (two hidden layers of width 256; line = binned medians). On the constructed mismatch there is essentially no relationship ( ρ=−0.05 ); under fitness extrapolation the trend inverts strongly ( ρ=−0.76 ), so sharper minima deploy better. This is a different quantity from the trace–gap correlation reported in § 4 ( ρ≈−0.52 ), which is measured on the full pool; the corresponding λmax –gap correlation here is −0.78 . Audited fixed-architecture pools; Cond7 is Lipophilicity (mismatch) and Hydro is Hydrophobic Core.
Proxy ∼λmax
Proxy ∼ gap
λmax∼ gap
condition
marg.
partial
marg.
partial
marg.
partial
Caco2 (random)
-0.13
+0.20
+0.13
+0.16
+0.08
+0.07
Caco2 (scaffold)
-0.14
+0.18
+0.13
+0.16
+0.02
-0.01
Lipo (random)
-0.39
-0.09
+0.21
+0.35
+0.07
-0.01
Lipo (scaffold)
-0.39
-0.09
+0.15
+0.23
+0.04
-0.01
Lipo (mismatch)
-0.40
-0.12
+0.15
+0.23
+0.04
-0.02
Appendix
Table 19: The proxy against exact curvature and against the generalisation gap. “marg.” is the marginal Spearman correlation across the sweep; “partial” additionally residualises on learning rate, width and depth. The learning rate is the confounder: it raises the proxy while lowering λmax , manufacturing the marginal association. Audited pools.
quantity
Caco2 (scaffold)
Lipo (mismatch)
Hydrophobic Core
λmax
+0.01
-0.05
-0.48
tr(H)
-0.02
-0.17
-0.52
∥w∥2tr(H)
+0.24
+0.33
-0.41
∥w∥2 alone
+0.18
+0.31
-0.21
Proxy (ours)
+0.17
+0.29
-0.31
tr(H) given ∥w∥2
+0.08
+0.07
-0.49
Appendix
Table 20: Hessian trace and approximate relative flatness, ∥w∥2tr(H) , on the three conditions where the trace was measured (720 networks trained, 642 after the audit; Hutchinson with 30 Rademacher probes). All entries are partial Spearman correlations with the generalisation gap given learning rate, width and depth. The final row shows that the trace adds almost nothing once the weight norm is controlled.
ablation
condition
reference
ref.
Proxy
λmax
Random
fixed arch.
Lipo (mismatch)
Val
0.510
0.512
0.543
0.533
fixed arch.
Hydrophobic Core
Val
23.245
22.565
25.532
22.654
validation-free
Caco2 (scaffold)
Train
0.476
0.459
0.482
0.496
validation-free
Lipo (scaffold)
Train
0.543
0.507
0.528
0.540
validation-free
Amylase
Train
3.182
3.015
3.117
2.847
validation-free
Hydrophobic Core
Train
22.871
22.837
23.121
23.272
Appendix
Table 21: Ablations, deployment MSE (median over replications, audited pools). Fixed architecture pins the network to two hidden layers of width 256; its Proxy and λmax columns select on the signal alone. Validation-free folds the validation split back into training, so only train-side selectors are defined, Train replaces Val as the reference, and the Proxy and λmax columns are rank sums with training MSE (Train+Proxy and Train+ λmax ).
condition
ρval
Val
Proxy
Random
Caco2 α=0.00
+0.52
0.354
0.342
0.367
Caco2 α=0.25
+0.60
0.339
0.347
0.380
Caco2 α=0.50
+0.51
0.451
0.401
0.448
Caco2 α=0.75
+0.43
0.710
0.668
0.792
Caco2 α=1.00
+0.40
0.834
0.755
0.913
Lipo α=0.00
+0.74
0.400
0.407
0.435
Appendix
Table 22: Synthetic shift-severity sweep. α=0 is an IID split and α=1 is pure feature extrapolation; validation is held in-distribution throughout. ρval degrades with severity, but no threshold appears at which the proxy reliably overtakes validation. Audited pools.
Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, and proteins. These models are built through post-training, yet how each stage shapes reasoning and generalization remains poorly understood. We study when post-training improves performance and when it induces over-specialization. Across genomics, transcriptomics, and proteins, we train and evaluate more than 100 biological reasoning models under controlled variation in backbone, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL), measuring both in-domain (ID) and out-of-domain (OOD) performance. We find that each post-training stage reshapes generalization in a distinct way rather than contributing uniform gains. CPT improves downstream performance by aligning models with biological language. SFT consistently increases ID performance but causes OOD performance to peak early and decline as models fit the training distribution. RL, when applied to strong SFT checkpoints with aligned rewards, improves OOD performance and partially recovers generalization. These results show that biological reasoning does not improve monotonically with additional supervision or compute. Instead, performance depends on how training stages are composed. Under fixed post-training budgets, the strongest ID-OOD trade-off comes from brief SFT, larger RL allocations, and asymmetric adaptation capacity across stages.
Lukas Fesser, Hanlin Zhang, Michelle M. Li +5
Harvard University · Google DeepMind · Google Research
Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which training recipe to apply. Making these decisions well requires reliable performance forecasts, yet the two commonly used signals are fundamentally limited. Cross-entropy loss is poorly aligned with downstream capabilities, and direct downstream evaluation is expensive, sparse, and often uninformative at early training stages. Instead, we propose to construct proxy metrics by aggregating token-level statistics, such as entropy, top-k accuracy, and expert token rank, from a candidate model's next token distribution over expert-written solutions. Across three settings, our proxies consistently outperform loss- and compute-based baselines: 1) For cross-family model selection, they rank a heterogeneous population of reasoning models with mean Spearman Rho = 0.81 (vs. Rho = 0.36 for cross-entropy loss); 2) For pretraining data selection, they reliably rank 25 candidate corpora for a target model at roughly 10,000× less compute than direct evaluation, pushing the Pareto frontier beyond existing methods; and 3) for training-time forecasting, they extrapolate downstream accuracy across an 18× compute horizon with roughly half the error of existing alternatives. Together, these results suggest that expert trajectories are a broadly useful source of signal for assessing model capabilities, enabling reliable performance forecasting throughout the model development life cycle.
Arkil Patel, Siva Reddy, Marius Mosbach +1
Mila – Quebec AI Institute & McGill University Canada · ServiceNow Research · Periodic Labs
A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release. But do we need to run every eval? We compile a public score matrix of 84 frontier models on 133 benchmarks (2,604 cells, 23.3% filled) and find it is approximately rank-2: a model's scores across all 133 benchmarks are largely determined by just two numbers. We confirm this in two ways: scores hidden from the matrix are best recovered using two factors, and two factors already explain over 90% of the variation among models on the benchmarks they share. Building on this, we design BenchPress: a logit-space rank-2 matrix completion method that recovers held-out scores to within 4.6 points, and a confidence layer that says when each prediction can be trusted. Using BenchPress, we find a subset of five benchmarks {GPQA-D, HLE, Codeforces, MMLU-Pro, ARC-AGI-1} that can recover the rest of a model's public scorecard to within 3.93 points. For a tighter inference budget, a cheaper set {GPQA-D, MMLU-Pro, Aider Polyglot, MATH-500, AIME 2026} can predict a model's evals to within 4.55. We release the score matrix, the BenchPress code, and an interactive tool that predicts any model's score on any benchmark.