Laboratories often face a new molecular assay with 16-64 labels and a bank of predictors whose training data and parameters are unavailable. The practical question is which frozen outputs to include in a small local model. AssayRouter treats completed assays as pseudo-targets and labels each candidate by its post-fit utility: the reduction in held-out discovery loss when the candidate is added to the local target predictor. A shared regressor learns to predict this utility from candidate behavior on the support set, without source identity; on a new assay, one frozen ranking selects four sources and separate labels fit a convex combiner. We train only on completed ChEMBL-MT assays and evaluate 24 external regression assays across six frozen interface families. AssayRouter-C lowers strict four-call negative log-likelihood (NLL) by 0.0409 relative to Support-CV@4. Frozen candidate-label permutations confirm that candidate-utility correspondence carries the transferred information, and leave-one-interface-out training shows that the mapping generalizes to unseen predictor families. Completed assays therefore provide transferable supervision for scarce-label routing through frozen prediction interfaces.
Figures & tables
Figure 1: Completed assays supervise frozen predictor routing. (1) A historical pseudo-target measures the reduction in discovery loss from adding one candidate. (2) A shared AssayRouter prior maps candidate behavior to utility without source identity. (3) On a new assay, the frozen prior scores the bank once, selects four sources, and fits simplex weights for the target before confirmation.
NLL by target labels
Overall
Method
16
32
64
NLL ↓
MAE ↓
RMSE ↓
Pred. ρ↑
AR-Marginal
2.2711
2.1052
1.9886
2.1216
23.8597
35.3358
0.3191
Support-CV
2.3287
2.1309
1.9938
2.1511
26.2446
38.0759
0.3033
Support-greedy
2.3806
2.1695
2.0410
2.1970
28.6550
40.3969
0.2949
Source-quality
2.4467
2.2451
2.0806
2.2575
29.3549
41.0480
0.3111
Residual
2.5339
2.2696
2.0806
2.2947
32.3550
44.8218
0.3053
Table 1: Strict four-call deployment. C, Δ , and Z denote centered, uncentered, and standardized standalone utility; Marginal denotes marginal-over-states utility. Every row uses one-shot constituent scoring; bold and underline mark the two best point estimates.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Selector
History
Full-bank query
Query-local
AssayRouter
Utility
No
No
Support-CV
None
No
No
Frozen-DES
None
Yes
Yes
MINE-WS ( Moura et al., 2021 )
Competence
Yes
Yes
All-source
None
Yes
No
Appendix
Table 2: Information available to the principal selectors. Query-local methods may change the selected set across confirmation molecules.
Alternative explanation
Matched test
Changed axis
No value from bank reuse
Target-only
Bank reuse
Generic source strength suffices
Source-quality
Utility supervision
Metadata/support statistics suffice
Size / SNR
Ranking signal
Profile compatibility suffices
Residual (strict)
Ranking signal
Candidate–label pairing is arbitrary
Frozen permutations
Ranking signal
Target-side search suffices
Support-CV@4
Access/timing
Appendix
Table 3: Coverage of alternative explanations relevant to the decision across five contract axes. Each control changes one scientific axis while preserving the frozen prediction contract wherever that axis permits.
Test
NLL effect
Positive breadth
Candidate-label intervention
0.1352
90.7% of cells
Unseen-interface intervention
0.0497
6/6 held-out families
Support-CV@4 vs. C
0.0409
77.8% of target–interface units
Residual compatibility vs. C
0.1844
83.3% of targets
Appendix
Table 4: Evidence-chain index. Effects favor AssayRouter-C; the final column reports the scientific unit appropriate to each test.
Comparison
Effect
Value
Direct intervention
Permuted − genuine NLL
+0.0463
Strict intervention
Permuted − genuine NLL
+0.0816
Unseen interface
Permuted − genuine NLL
+0.0299
Unseen interfaces
Positive held-out families
5/6
Strict search
Support-CV − genuine NLL
+0.0439
Appendix
Table 5: Support-scale Qdisc sensitivity. Values are paired point effects; all five prespecified criteria are met.
Method
NLL ↓
MAE ↓
RMSE ↓
Spearman ↑
AssayRouter- Δ
2.1142
23.5806
35.3864
0.2961
AssayRouter-C
2.1102
23.3872
35.0709
0.3011
AssayRouter-Z
2.1181
23.6652
35.3589
0.3124
Post-fit loss
2.1684
25.9818
37.2848
0.3002
Post-fit rank
2.1515
24.9610
36.3301
0.3228
Appendix
Table 6: Strict matched post-fit controls. Point estimates are averaged after each four-call constituent is scored.
Estimand
NLL effect
95% interval
Cell win rate
Cross-fit
0.0814
[0.0427,0.1246]
0.7222
Strict
0.1352
[0.0852,0.1915]
0.9074
Appendix
Table 7: Direct AssayRouter-C candidate-label intervention. NLL effects are permuted minus genuine; win rate is the fraction of 432 target–interface–budget cells favoring the genuine mapping. Intervals use target–interface clustered resampling.
Role
Collection
Targets
Interfaces
Historical supervision
ChEMBL-MT ( Adrian et al., 2025 )
20
6
Temporal holdout
ExpansionRx ( OpenADMET Consortium, 2026 )
9
6
Sparse-panel holdout
Biogen ( Fang et al., 2023 )
6
6
Scaffold holdout
TDC regression ( Huang et al., 2021 )
9
6
Appendix
Table 8: Training and frozen external evaluation. External target identities and confirmation labels are absent from router training and model selection.
Molecular representation
Source predictor
Reuse mode
Morgan FP ( Rogers and Hahn, 2010 )
Ridge
Contract
RDKit2D ( Landrum and RDKit Contributors, n.d. )
LightGBM ( Ke et al., 2017 )
Contract
CheMeleon emb. ( Burns et al., 2025 )
Ridge
Frozen repr.
ChemBERTa2 emb. ( Ahmad et al., 2022 )
LightGBM ( Ke et al., 2017 )
Frozen repr.
Molecular graph
GIN ( Xu et al., 2019 )
Per-source
CheMeleon repr. ( Burns et al., 2025 )
Prediction head
Fine-tuned
Appendix
Table 9: The six formal interfaces to frozen predictors. Abbreviations keep each cell on one line: FP, emb., repr., and GIN denote fingerprint, embedding, representation, and graph-isomorphism network. A bank contains independently trained source-assay predictors from one interface family.
Scope
Mean difference
Overall
−0.0038
16 labels
−0.0067
32 labels
−0.0002
64 labels
−0.0043
Biogen
−0.0035
ExpansionRx
0.0008
Appendix
Table 10: Matched supervision refinements. Differences are marginal-over-states utility NLL minus standalone-utility NLL; positive values favor the primary utility prior.
Router
Historical target
State features
External NLL
AssayRouter (standalone-utility prior)
Δ(s∣∅)
No
2.0418
Marginal-over-states utility
EAΔ(s∣A)
No
2.0380
State-conditioned extension
Δ(s∣A)
Yes
2.0392
Appendix
Table 11: Prediction refinement and external decision quality. Historical out-of-fold (OOF) mean-squared error (MSE) is defined only for the shared marginal-over-states label; lower is better.
Scope
Marginal utility MSE
State-conditioned MSE
Relative reduction
Overall
0.02064
0.01949
5.58%
16 labels
0.02584
0.02419
6.41%
32 labels
0.02061
0.01923
6.72%
64 labels
0.01546
0.01504
2.66%
Appendix
Table 12: Nested post-fit utility-prediction audit. Both columns use frozen target-grouped out-of-fold predictions for the same marginal labels.
Frozen mapping
Mean NLL effect
Cell win rate
1
0.0494
73.38%
2
0.0477
70.83%
3
0.0146
60.65%
4
0.0339
72.69%
5
0.0454
71.53%
Appendix
Table 13: Independent candidate-label interventions under the cross-fit prediction-average estimand. Each row uses a distinct frozen bijection; effects are permuted minus genuine NLL.
Table 15: Stability-focused cross-fit estimand. Absolute support-standardized robust NLL on 24 external targets excluded from router fitting and six prediction interfaces; lower is better. All sparse constituent models have four-source combiner fan-in, but their information permissions differ as described in the text.
Scope
Mean difference
Win rate
Overall
0.0494
73.38%
16 labels
0.0813
84.03%
32 labels
0.0415
72.22%
64 labels
0.0254
63.89%
Biogen
0.0030
50.00%
ExpansionRx
0.0865
85.19%
Appendix
Table 16: Direct standalone-utility permutation. Differences are permuted-router NLL minus genuine AssayRouter NLL; positive values favor genuine post-fit labels.
Metric
Mean effect
Robust NLL
0.0494
MAE
3.5805
RMSE
3.1850
Spearman correlation
0.0327
Appendix
Table 17: Secondary metrics for the direct standalone-utility permutation. NLL, MAE, and RMSE are permuted minus genuine; Spearman is genuine minus permuted.
Scope
Mean difference
Overall
0.0870
16 labels
0.1260
32 labels
0.0816
64 labels
0.0533
Biogen
0.0263
ExpansionRx
0.1081
Appendix
Table 18: Strict four-call standalone-utility intervention. Differences are permuted minus genuine NLL; positive values favor AssayRouter.
Metric
Mean effect
Robust NLL
0.0870
MAE
4.6262
RMSE
4.9810
Spearman correlation
0.0405
Appendix
Table 19: Overall secondary effects for the strict four-call intervention. Losses are permuted minus genuine; Spearman is genuine minus permuted.
Held-out scope
Permuted
Support-CV
Overall
0.0497
0.1116
16 labels
0.0697
0.1769
32 labels
0.0492
0.1029
64 labels
0.0301
0.0550
Biogen
−0.0023
0.0367
ExpansionRx
0.0770
0.1594
Appendix
Table 20: AssayRouter-C leave-one-interface-out evaluation. Differences are comparator minus genuine-router NLL after the named family is removed from historical supervision; positive values favor AssayRouter-C.
Router
Scope
Targets
Retention
Mean difference
Direct
Full
9
100.00%
0.0433
Direct
H -clean
9
62.38%
0.0419
Direct
S -clean
8
55.97%
0.0403
Direct
Union-clean
8
42.71%
0.0385
LOIO
Full
9
100.00%
0.0525
LOIO
H -clean
9
62.38%
0.0512
Appendix
Table 21: TDC exact-parent nonoverlap sensitivity. Differences are permuted minus genuine NLL for the AssayRouter- Δ replication; positive values favor historical candidate–utility correspondence. Retention is the mean confirmation fraction.
Scope
Support-CV NLL
Difference
Overall
2.0448
0.0030
16 labels
2.1806
0.0200
32 labels
2.0315
0.0015
64 labels
1.9223
−0.0125
Appendix
Table 22: Matched Support-CV@4 cross-fit comparison. NLL is computed after averaging the same 16 directional prediction vectors. Differences are Support-CV minus AssayRouter- Δ .
Scope
Mean difference
Overall
0.0409
16 labels
0.0736
32 labels
0.0413
64 labels
0.0078
Biogen
0.0005
ExpansionRx
0.0716
Appendix
Table 23: Strict four-call Support-CV@4 comparison. Differences are Support-CV NLL minus AssayRouter-C NLL; positive values favor the frozen historical prior.
Quantity
AssayRouter-C
Support-CV@4
Historical label NNLS fits
28,800
0
Selection NNLS fits per direction
0
246.5
Selection NNLS fits, full evaluation
0
20,445,696
Strict confirmation calls per constituent
4
4
Constituents per cross-fit episode
16
16
Appendix
Table 24: Comparison scopes. Selection fit counts measure subset construction on the target; strict confirmation calls define the accuracy budget. Counts exclude total wall-clock latency.
Comparator
Biogen
ExpansionRx
TDC
50 random four-source sets
−0.0026
0.0713
0.0275
Support-greedy selection ( Caruana et al., 2004 )
−0.0047
0.0547
0.0356
Frozen-DES@4
−0.0055
0.0013
−0.0044
MINE-WS@4 ( Moura et al., 2021 )
−0.0010
0.0421
0.0229
All-source convex
−0.0019
0.0440
−0.0274
Appendix
Table 25: Cross-fit greedy and information-rich diagnostics by collection. Comparator NLL minus AssayRouter- Δ NLL, averaged over interfaces and support budgets; positive values favor the post-fit utility prior.
Interface
Random
Greedy ( Caruana et al., 2004 )
DES@4
All-source
Morgan / Ridge
0.0333
0.0292
−0.0010
0.0078
RDKit2D / LGBM
0.0445
0.0413
−0.0005
0.0140
CheMeleon / Ridge
0.0311
0.0304
−0.0096
−0.0023
ChemBERTa2 / LGBM
0.0474
0.0459
0.0070
0.0100
GIN
0.0367
0.0351
0.0026
0.0126
CheMeleon / FT
0.0254
0.0141
−0.0138
−0.0078
Appendix
Table 26: Cross-fit greedy and information-rich diagnostics by interface. Comparator NLL minus AssayRouter- Δ NLL, averaged over external targets and budgets; positive values favor the post-fit utility prior.
Comparator and scope
Mean difference
State-conditioned extension, overall
0.0011
State-conditioned, 16 labels
0.0051
State-conditioned, 32 labels
−0.0012
State-conditioned, 64 labels
−0.0004
Marginal labels permuted, overall
0.0532
Marginal labels permuted, 16 labels
0.0824
Appendix
Table 27: Matched marginal-over-states mechanism tests. Differences are comparator NLL minus genuine marginal-over-states NLL; positive values favor the genuine labels.
Method
Runs
Macro Pearson r
Random Forest ( Breiman, 2001 )
6
0.6683
LightGBM ( Ke et al., 2017 )
6
0.7069
MolSetRep GINE ( Boulougouri et al., 2024 )
18
0.6373
MolSetRep SR-GINE ( Boulougouri et al., 2024 )
18
0.7259
Appendix
Table 28: Official full-training reproduction on all six public Biogen endpoints. Pearson r is macro-averaged across endpoints.
Protocol
Biogen (4)
ExpansionRx (9)
Author KERMT task-specific
0.3320
0.3750
Author best Contrastive KERMT
0.3210
0.3590
Official-checkpoint reproduction
0.3515
0.2836
Appendix
Table 29: Full-training KERMT comparison (macro MAE; lower is better). Author values are transcribed from the official manuscript source; reproductions use the released Contrastive KERMT checkpoint, five seeds, and the frozen public splits.
School of Computer Science, Peking University · Key Laboratory of High Confidence Software Technologies, Peking University, Ministry of Education · School of Computer Science, Wuhan University +1
L3S Research Center, Leibniz Universität Hannover, Appelstraße 9a, 30167 Hannover, Germany · Institute of Data Science (Knowledge-Based Systems), Faculty of Electrical Engineering and Computer Science,2026 Leibniz Universität Hannover, Appelstraße 9a, 30167 Hannover, Germany
Molecular AI, Discovery Sciences R&D, AstraZeneca · Department of Information Technology, Uppsala University · Robotics, Perception & Learning, KTH Royal Institute of Technology +3