The Mechanics of Delta Learning: Target Design for Generalizable Scientific Machine Learning
Authors: Kareem M. Gameel, Ihor Neporozhnii, Sjoerd Hoogland, Oleksandr Voznyy
Organizations: Department of Physical and Environmental Sciences, University of Toronto Scarborough Alliance for AI-Accelerated Materials Discovery (A3MD) Toronto, Ontario, Canada · Alliance for AI-Accelerated Materials Discovery (A3MD) Toronto, Ontario, Canada · Department of Chemistry, University of Toronto Department of Physical and Environmental Sciences, University of Toronto Scarborough Alliance for AI-Accelerated Materials Discovery (A3MD) Toronto, Ontario, Canada
In scientific machine learning, Δ-learning trains models on residual errors relative to physical baselines, assuming that more accurate baselines with smaller residual scales inherently improve downstream performance. Here, we demonstrate that residual scale alone is an insufficient heuristic for learnability. Evaluating molecular graph neural networks on total energy targets, we show that complex local descriptor baselines can yield small residual targets that are disproportionately rough within architecture-informed proxy spaces and harder to learn relative to their scale. Conversely, semi-empirical baseline reduces both scale and normalized roughness, improving in-domain and out-of-domain prediction. We introduce scale-normalized graph Dirichlet roughness (DIQR) as a pre-training diagnostic for residual learnability and establish baseline complementarity as a core target-design principle, elevating target space formulation alongside model architecture as a key axis for scientific machine learning.
Figures & tables
Baseline model
Input features
Fitting approach
Physical/representation target
BoB
Typed-bond counts ( RDKit )
Linear / MLP
Compositional and functional-group variation
PRE
Pairwise distances ( rij )
Linear / MLP
Local two-body geometry (SchNet proxy)
PRE+A
Distances and angles ( rij,θijk )
Linear / MLP
Local three-body geometry (GemNet proxy)
xTB
Atomic identities and molecular coordinates
Semiempirical QM
Non-local electronic and higher-order interactions
DFTB †
Atomic identities and molecular coordinates
Semiempirical QM
Molecule-wide electronic structure; supporting comparison with xTB
Table 1 : Baseline models used for target transformation.
Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular domain, where minor changes can drastically alter properties. In this work, we show that semi-supervised methods that rely on an ensemble consensus can boost predictive accuracy across a diverse range of molecular datasets, task types, and graph neural network architectures. We find that training with an ensemble consensus objective increases robustness in models and exhibits an effect similar to knowledge distillation; an individual member of an ensemble trained this way outperforms a full ensemble trained in a traditional supervised fashion in almost all cases. In addition, this type of semi-supervised training reduces calibration error.
Scientific foundation models (SciFMs) aim to learn generalizable representations of physical systems governed by partial differential equations (PDEs), enabling transfer across tasks and domains. While physics-informed methods, which leverage PDE residuals as supervisory signals, have shown promise in scientific machine learning (SciML) for improving accuracy and reducing data requirements, their potential in the context of SciFMs remains relatively unexplored. In this evaluation study, we investigate whether (and how) physics-informed pre-training improves the generalization, robustness, and data efficiency of SciFMs. We conduct systematic experiments across a diverse set of PDEs, ranging from simple problems with periodic boundary conditions to more challenging systems such as the Navier-Stokes equations and non-periodic geometries. Our results show that physics-informed pre-training provides clear benefits in nice,'' e.g., structured, well-aligned settings: it enhances generalization and reduces data dependence, compared to data-only pre-training. However, these advantages diminish significantly as the downstream tasks become harder,'' e.g., as they involve discontinuities or deviate from the pre-training distribution. In complex or structurally different problems, such as those involving new boundary conditions or PDE operators, physics-informed models may perform only on par with---or even worse---than data-driven baselines. While residual-based pre-training helps in idealized regimes, realizing broadly transferable SciFMs will likely require subtler spatiotemporal inductive biases and more principled integration of physical knowledge into model architectures.
Serge Kotchourko, Amin Totounferoush, Michael W. Mahoney +1
University of Stuttgart, Germany · ICSI, LBNL, and University of California, Berkeley, USA
General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and performance required for drug discovery tasks. Simply increasing model size or introducing reasoning tokens does not yield significant performance gains. To address this gap, we introduce the MMAI Gym for Science, a one-stop shop molecular data formats and modalities as well as task-specific reasoning, training, and benchmarking recipes designed to teach foundation models the 'language of molecules' in order to solve practical drug discovery problems. We use MMAI Gym to train an efficient Liquid Foundation Model (LFM) for these applications, demonstrating that smaller, purpose-trained foundation models can outperform substantially larger general-purpose or specialist models on molecular benchmarks. Across essential drug discovery tasks - including molecular optimization, ADMET property prediction, retrosynthesis, drug-target activity prediction, and functional group reasoning - the resulting model achieves near specialist-level performance and, in the majority of settings, surpasses larger models, while remaining more efficient and broadly applicable in the domain.
Maksim Kuznetsov, Zulfat Miftahutdinov, Rim Shayakhmetov +17