AssayRouter: Historical Utility Priors for Frozen Molecular Predictor Routing
Organizations: School of Artificial Intelligence, Shenzhen University · EasternDawn · School of Computer Science, University of Nottingham Ningbo
Abstract
Laboratories often face a new molecular assay with 16-64 labels and a bank of predictors whose training data and parameters are unavailable. The practical question is which frozen outputs to include in a small local model. AssayRouter treats completed assays as pseudo-targets and labels each candidate by its post-fit utility: the reduction in held-out discovery loss when the candidate is added to the local target predictor. A shared regressor learns to predict this utility from candidate behavior on the support set, without source identity; on a new assay, one frozen ranking selects four sources and separate labels fit a convex combiner. We train only on completed ChEMBL-MT assays and evaluate 24 external regression assays across six frozen interface families. AssayRouter-C lowers strict four-call negative log-likelihood (NLL) by 0.0409 relative to Support-CV@4. Frozen candidate-label permutations confirm that candidate-utility correspondence carries the transferred information, and leave-one-interface-out training shows that the mapping generalizes to unseen predictor families. Completed assays therefore provide transferable supervision for scarce-label routing through frozen prediction interfaces.
Figures & tables
| NLL by target labels | Overall | ||||||
|---|---|---|---|---|---|---|---|
| Method | 16 | 32 | 64 | NLL | MAE | RMSE | Pred. |
| AR-Marginal | 2.2711 | 2.1052 | 1.9886 | 2.1216 | 23.8597 | 35.3358 | 0.3191 |
| Support-CV | 2.3287 | 2.1309 | 1.9938 | 2.1511 | 26.2446 | 38.0759 | 0.3033 |
| Support-greedy | 2.3806 | 2.1695 | 2.0410 | 2.1970 | 28.6550 | 40.3969 | 0.2949 |
| Source-quality | 2.4467 | 2.2451 | 2.0806 | 2.2575 | 29.3549 | 41.0480 | 0.3111 |
| Residual | 2.5339 | 2.2696 | 2.0806 | 2.2947 | 32.3550 | 44.8218 | 0.3053 |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Selector | History | Full-bank query | Query-local |
|---|---|---|---|
| AssayRouter | Utility | No | No |
| Support-CV | None | No | No |
| Frozen-DES | None | Yes | Yes |
| MINE-WS ( Moura et al., 2021 ) | Competence | Yes | Yes |
| All-source | None | Yes | No |
| Alternative explanation | Matched test | Changed axis |
|---|---|---|
| No value from bank reuse | Target-only | Bank reuse |
| Generic source strength suffices | Source-quality | Utility supervision |
| Metadata/support statistics suffice | Size / SNR | Ranking signal |
| Profile compatibility suffices | Residual (strict) | Ranking signal |
| Candidate–label pairing is arbitrary | Frozen permutations | Ranking signal |
| Target-side search suffices | Support-CV@4 | Access/timing |
| Test | NLL effect | Positive breadth |
|---|---|---|
| Candidate-label intervention | 0.1352 | 90.7% of cells |
| Unseen-interface intervention | 0.0497 | 6/6 held-out families |
| Support-CV@4 vs. C | 0.0409 | 77.8% of target–interface units |
| Residual compatibility vs. C | 0.1844 | 83.3% of targets |
| Comparison | Effect | Value |
|---|---|---|
| Direct intervention | Permuted genuine NLL | |
| Strict intervention | Permuted genuine NLL | |
| Unseen interface | Permuted genuine NLL | |
| Unseen interfaces | Positive held-out families | |
| Strict search | Support-CV genuine NLL |
| Method | NLL | MAE | RMSE | Spearman |
|---|---|---|---|---|
| AssayRouter- | 2.1142 | 23.5806 | 35.3864 | 0.2961 |
| AssayRouter-C | 2.1102 | 23.3872 | 35.0709 | 0.3011 |
| AssayRouter-Z | 2.1181 | 23.6652 | 35.3589 | 0.3124 |
| Post-fit loss | 2.1684 | 25.9818 | 37.2848 | 0.3002 |
| Post-fit rank | 2.1515 | 24.9610 | 36.3301 | 0.3228 |
| Estimand | NLL effect | 95% interval | Cell win rate |
|---|---|---|---|
| Cross-fit | 0.0814 | 0.7222 | |
| Strict | 0.1352 | 0.9074 |
| Role | Collection | Targets | Interfaces |
|---|---|---|---|
| Historical supervision | ChEMBL-MT ( Adrian et al., 2025 ) | 20 | 6 |
| Temporal holdout | ExpansionRx ( OpenADMET Consortium, 2026 ) | 9 | 6 |
| Sparse-panel holdout | Biogen ( Fang et al., 2023 ) | 6 | 6 |
| Scaffold holdout | TDC regression ( Huang et al., 2021 ) | 9 | 6 |
| Molecular representation | Source predictor | Reuse mode |
|---|---|---|
| Morgan FP ( Rogers and Hahn, 2010 ) | Ridge | Contract |
| RDKit2D ( Landrum and RDKit Contributors, n.d. ) | LightGBM ( Ke et al., 2017 ) | Contract |
| CheMeleon emb. ( Burns et al., 2025 ) | Ridge | Frozen repr. |
| ChemBERTa2 emb. ( Ahmad et al., 2022 ) | LightGBM ( Ke et al., 2017 ) | Frozen repr. |
| Molecular graph | GIN ( Xu et al., 2019 ) | Per-source |
| CheMeleon repr. ( Burns et al., 2025 ) | Prediction head | Fine-tuned |
| Scope | Mean difference |
|---|---|
| Overall | |
| 16 labels | |
| 32 labels | |
| 64 labels | |
| Biogen | |
| ExpansionRx | 0.0008 |
| Router | Historical target | State features | External NLL |
|---|---|---|---|
| AssayRouter (standalone-utility prior) | No | 2.0418 | |
| Marginal-over-states utility | No | 2.0380 | |
| State-conditioned extension | Yes | 2.0392 |
| Scope | Marginal utility MSE | State-conditioned MSE | Relative reduction |
|---|---|---|---|
| Overall | 0.02064 | 0.01949 | 5.58% |
| 16 labels | 0.02584 | 0.02419 | 6.41% |
| 32 labels | 0.02061 | 0.01923 | 6.72% |
| 64 labels | 0.01546 | 0.01504 | 2.66% |
| Frozen mapping | Mean NLL effect | Cell win rate |
|---|---|---|
| 1 | 0.0494 | 73.38% |
| 2 | 0.0477 | 70.83% |
| 3 | 0.0146 | 60.65% |
| 4 | 0.0339 | 72.69% |
| 5 | 0.0454 | 71.53% |
| Alternative | Cross-fit | Strict |
|---|---|---|
| Target-only | 0.2403 | 0.2197 |
| Source-quality | 0.0969 | 0.1433 |
| Train-size | 0.0912 | 0.1174 |
| Source-SNR | 0.0789 | 0.1304 |
| Support-CV | 0.0030 | 0.0369 |
| Method | 16 labels | 32 labels | 64 labels |
|---|---|---|---|
| AssayRouter (standalone-utility prior) | 2.1606 | 2.0300 | 1.9348 |
| Matched refinements | |||
| Marginal-over-states utility | 2.1539 | 2.0297 | 1.9305 |
| State-conditioned extension | 2.1589 | 2.0285 | 1.9301 |
| Search and information-rich controls | |||
| Support-CV@4 ensemble | 2.1806 | 2.0315 | 1.9223 |
| Scope | Mean difference | Win rate |
|---|---|---|
| Overall | 0.0494 | 73.38% |
| 16 labels | 0.0813 | 84.03% |
| 32 labels | 0.0415 | 72.22% |
| 64 labels | 0.0254 | 63.89% |
| Biogen | 0.0030 | 50.00% |
| ExpansionRx | 0.0865 | 85.19% |
| Metric | Mean effect |
|---|---|
| Robust NLL | 0.0494 |
| MAE | 3.5805 |
| RMSE | 3.1850 |
| Spearman correlation | 0.0327 |
| Scope | Mean difference |
|---|---|
| Overall | 0.0870 |
| 16 labels | 0.1260 |
| 32 labels | 0.0816 |
| 64 labels | 0.0533 |
| Biogen | 0.0263 |
| ExpansionRx | 0.1081 |
| Metric | Mean effect |
|---|---|
| Robust NLL | 0.0870 |
| MAE | 4.6262 |
| RMSE | 4.9810 |
| Spearman correlation | 0.0405 |
| Held-out scope | Permuted | Support-CV |
|---|---|---|
| Overall | 0.0497 | 0.1116 |
| 16 labels | 0.0697 | 0.1769 |
| 32 labels | 0.0492 | 0.1029 |
| 64 labels | 0.0301 | 0.0550 |
| Biogen | 0.0367 | |
| ExpansionRx | 0.0770 | 0.1594 |
| Router | Scope | Targets | Retention | Mean difference |
|---|---|---|---|---|
| Direct | Full | 9 | 100.00% | 0.0433 |
| Direct | -clean | 9 | 62.38% | 0.0419 |
| Direct | -clean | 8 | 55.97% | 0.0403 |
| Direct | Union-clean | 8 | 42.71% | 0.0385 |
| LOIO | Full | 9 | 100.00% | 0.0525 |
| LOIO | -clean | 9 | 62.38% | 0.0512 |
| Scope | Support-CV NLL | Difference |
|---|---|---|
| Overall | 2.0448 | 0.0030 |
| 16 labels | 2.1806 | 0.0200 |
| 32 labels | 2.0315 | 0.0015 |
| 64 labels | 1.9223 |
| Scope | Mean difference |
|---|---|
| Overall | 0.0409 |
| 16 labels | 0.0736 |
| 32 labels | 0.0413 |
| 64 labels | 0.0078 |
| Biogen | 0.0005 |
| ExpansionRx | 0.0716 |
| Quantity | AssayRouter-C | Support-CV@4 |
|---|---|---|
| Historical label NNLS fits | 28,800 | 0 |
| Selection NNLS fits per direction | 0 | 246.5 |
| Selection NNLS fits, full evaluation | 0 | 20,445,696 |
| Strict confirmation calls per constituent | 4 | 4 |
| Constituents per cross-fit episode | 16 | 16 |
| Comparator | Biogen | ExpansionRx | TDC |
|---|---|---|---|
| 50 random four-source sets | 0.0713 | 0.0275 | |
| Support-greedy selection ( Caruana et al., 2004 ) | 0.0547 | 0.0356 | |
| Frozen-DES@4 | 0.0013 | ||
| MINE-WS@4 ( Moura et al., 2021 ) | 0.0421 | 0.0229 | |
| All-source convex | 0.0440 |
| Interface | Random | Greedy ( Caruana et al., 2004 ) | DES@4 | All-source |
|---|---|---|---|---|
| Morgan / Ridge | 0.0333 | 0.0292 | 0.0078 | |
| RDKit2D / LGBM | 0.0445 | 0.0413 | 0.0140 | |
| CheMeleon / Ridge | 0.0311 | 0.0304 | ||
| ChemBERTa2 / LGBM | 0.0474 | 0.0459 | 0.0070 | 0.0100 |
| GIN | 0.0367 | 0.0351 | 0.0026 | 0.0126 |
| CheMeleon / FT | 0.0254 | 0.0141 |
| Comparator and scope | Mean difference |
| State-conditioned extension, overall | 0.0011 |
| State-conditioned, 16 labels | 0.0051 |
| State-conditioned, 32 labels | |
| State-conditioned, 64 labels | |
| Marginal labels permuted, overall | 0.0532 |
| Marginal labels permuted, 16 labels | 0.0824 |
| Method | Runs | Macro Pearson |
|---|---|---|
| Random Forest ( Breiman, 2001 ) | 6 | 0.6683 |
| LightGBM ( Ke et al., 2017 ) | 6 | 0.7069 |
| MolSetRep GINE ( Boulougouri et al., 2024 ) | 18 | 0.6373 |
| MolSetRep SR-GINE ( Boulougouri et al., 2024 ) | 18 | 0.7259 |
| Protocol | Biogen (4) | ExpansionRx (9) |
|---|---|---|
| Author KERMT task-specific | 0.3320 | 0.3750 |
| Author best Contrastive KERMT | 0.3210 | 0.3590 |
| Official-checkpoint reproduction | 0.3515 | 0.2836 |