cs.LGSep 30, 2026

A Generalisation Signal Need Not Be a Model-Selection Signal

Authors: Aditya Nagarsekar, M P Ashish Bhat, Aadi Nesarkar, Vrishti Godhwani, Rahul Yedida, Aditya Challa, Danda Sravan, Snehanshu Saha

Organizations: Department of CS&IS, BITS Pilani, K K Birla Goa Campus · LexisNexis Legal & Professional · Center for AI and Supercomputing, Mahindra University

Abstract

Model selection in computational biology often relies on validation data drawn from the training regime, even when deployment lies outside it. When validation no longer preserves which model is best, a natural alternative is to rank candidates using properties of the trained network itself. We test this idea using a novel, forward-only proxy motivated by the norm of the Hessian, alongside common Hessian measures, across molecular property, protein fitness, and drug-response tasks. Contrary to our hypothesis, geometry does not become more useful as validation Spearman correlation deteriorates: augmenting validation helps some shifts but significantly harms others. More surprisingly, the proxy still correlates with generalisation gap on most tasks even when Hessian trace and top-eigenvalue relationships are weak or reversed, yet this signal does not reliably identify the deployment-best model. A curvature bound need not preserve cross-model rankings, and low geometric scores can even favour collapsed predictors. Thus, a generalisation signal need not be a model-selection signal.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Post-Training Shapes Biological Reasoning Models

    Jun 15, 2026Lukas Fesser, Hanlin Zhang, Michelle M. Li +5Large Reasoning ModelsPost-Training

  2. Forecasting Downstream Performance of LLMs With Proxy Metrics

    May 18, 2026Arkil Patel, Siva Reddy, Marius Mosbach +1Cross-Entropy LossesToken-Level Uncertainty

  3. You Don't Need to Run Every Eval

    Jun 22, 2026Yuchen Zeng, Dimitris PapailiopoulosModel EvaluationFrontier Models