Langevin-Informed Transfer Learning: Replacing Target Samples by Black-Box Feedback
Authors: Vladimir R. Kostic, Karim Lounici, Hélène Halconruy, Timothée Devergne, Michele Parrinello, Massimiliano Pontil
Organizations: CSML, Istituto Italiano di Tecnologia · University of Novi Sad · CMAP-Ecole Polytechnique · SAMOVAR, Télécom Sud-Paris · MODAL’X, Université Paris Nanterre · CSML & ATSIM, Istituto Italiano di Tecnologia · ATSIM, Istituto Italiano di Tecnologia · AI Centre, University College London
Many scientific and machine learning systems, from molecular dynamics to diffusion models and beyond, are governed by stochastic dynamics with low-dimensional structure, evolving on slow timescales. However, target trajectories, used to identify and interpret such dynamics, are often inaccessible: only biased or static samples that explore the underlying manifold are available. We introduce Langevin-Informed Transfer Learning (LITL), a framework for recovering target Langevin dynamics from biased source samples using only black-box feedback. LITL learns the leading spectral structure of the target infinitesimal generator and the projected drift through Dirichlet representation learning, enabling kinetic reconstruction in spectral form and slow-manifold gradient field estimation. We further introduce a spherical variant well suited to steering normalized latent representations commonly used in learning systems toward desired objectives. We establish finite-sample guarantees for eigenvalue, eigenfunction, and projected drift estimation in Sobolev norms, thereby ensuring generalization of these quantities and their first-order derivatives. Empirically, LITL recovers physical transition timescales from biased molecular simulations, builds kinetic structure from static samples of generative models, reconstructs spherical symmetries of physical systems, and enables post-hoc latent steering of trained neural networks under black-box feedback. Together, these results position spectral operator learning as a practical framework for recovering stochastic dynamics under distribution shift and unlock applications across machine learning and the physical sciences.
Figures & tables
Method
λ1=0.22
λ2=15.37
λ3=16.06
λ4=47.31
Devergne et al. (2024)
0.621 ± 0.01
46.6 ± 0.3
49.1 ± 0.5
146 ± 2
Zhang et al. (2022)
1 ± 2
2 ± 2
2 ± 2
8 ± 7
LITL
0.23 ± 0.02
15.315 ± 0.01
16.09 ± 0.04
46.9 ± 0.1
Table 1: Eigenvalue estimation in the double well experiment.
Figure 1 : Free-energy landscapes learned from MD and BioEmu samples using LITL. State occupation probabilities computed from the recovered generators reveal kinetic mismatch between BioEmu and physical MD dynamics.
Figure 2 : Leading eigenfunctions of the Langevin generator for a cobalt nanoparticle potential in polar coordinate representation ψ(θ,φ) associated to eigenvalues ∣λ1∣≈1.36 and ∣λ2∣≈∣λ3∣≈2.5 . The first non-trivial eigenfunction ψ1 reveals the slowest transition pathway between metastable states, while couple (ψ2,ψ3) captures faster symmetric modes orthogonal to the first one.
Figure 3 : Accuracy–fairness Pareto curve for the Adult Income experiment. Each point corresponds to a different number of LITL steps. LITL hyperparameters are calibrated on audit set, and then used to achieve substantial reduction in TPR disparity with a limited decrease in accuracy on a test set.
Method
Rel. Acc. [%] ↑
Rel. TPR gap [%] ↑
LITL (full)
97.11 ± 1.00
94.16 ± 4.04
LITL (local proxy)
99.91 ± 0.09
8.02 ± 2.41
LITL (no feedback)
99.27 ± 0.37
59.88 ± 5.08
LITL (bad truncation)
95.68 ± 0.89
70.68 ± 5.72
Retraining baseline
101.32 ± 0.21
94.60 ± 4.30
Fine-tuning baseline
100.25 ± 0.17
90.11 ± 5.78
Table 2: Ablation study and baseline comparisons in the fairness experiment over 10 trials with recorded relative accuracy and TPR gap change w.r.t. the original classifier.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Learning
Transfer
Needs Target
Needs
Representative
Method
Mechanism
Samples or
Retraining
References
Gradients
Fine-tuning /
Parameters /
Yes
Yes/No
[ 60 ]
Feature Transfer
Features
[ 44 ]
Domain Adaptation /
Reweighted Loss
Yes/No
Yes
[ 3 ]
Covariate Shift
[ 52 ]
Appendix
Table 3: Comparison of LITL with related transfer learning, robustness, sampling, and diffusion-based approaches. Latent diffusion models transfer learned generative dynamics through retraining in latent spaces, whereas LITL transfers spectral structure of the underlying Langevin generator without requiring target gradients, retraining, or access to target labels.
Aspect
Diffusion / Score-Based Models
LITL (Ours)
Task
Generate samples from a given data distribution
Create dynamics from log density discrepancies
Observed data
Target samples or noisy marginals
Samples from a source Gibbs distribution
Learned object
Score field ∇logpt(x)
Generator eigenspaces & projected score field
Score estimation
Full, pointwise
Spectrally projected (slow manifold)
Time dependence
Time-dependent
Time-homogeneous
Underlying dynamics
Artificial noising process
Physical / optimization-induced Langevin dynamics
Appendix
Table 4: Compact comparison between diffusion/score-based generative models and Langevin-Informed Transfer Learning (LITL). Diffusion models learn scores to reproduce observed data distributions, whereas LITL transfers kinetic structure to construct unseen dynamics without access to target samples or gradients.
Property
Sample Splitting
U-Statistics
Computational complexity
O(bm2d)
O(b2md)
Handles trivial eigenpair
Yes
Requires centering
Batch efficiency
Better for large b
Better for large m
Appendix
Table 5 : Comparison of empirical loss estimation strategies.
Figure 4 : Top left panel: the different potentials involved in this experiment in units of 1/β . Other panels: comparison of the eigenfunctions obtained with LITL and with the methods presented in [ 17 ] and [ 62 ] . Error bars have been computed using 4 different models with different random initial parameters
Figure 5 : 1D double-well Langevin dynamics. Gradient of the target potential projected on the leading eigenspaces (left), and its estimation from data (right). Uncertainty is computed over 4 different models
Figure 6 : 1D double-well Langevin dynamics. Target flow from an initial distribution and its estimation from data.
Figure 7 : First triplet of leading eigenfunctions of the Langevin generator for a cobalt nanoparticle potential. First row show 3D visualization on the unit sphere S2 , colored by eigenfunction value, while the second one shows polar coordinate representation ψ(θ,φ) . Third row shows reult obtained by LITL. Multiplicity of learned eigenvalues (i.e. time-scales) and spatial symmetry of eigenfunction are consistent with symmetries of the potential energy, and eigenvalue and eignfunction errors are reported in the header.
Figure 8 : Second triplet of leading eigenfunctions of the Langevin generator for a cobalt nanoparticle potential. First row show 3D visualization on the unit sphere S2 , colored by eigenfunction value, while the second one shows polar coordinate representation ψ(θ,φ) . Third row shows reult obtained by LITL. Multiplicity of learned eigenvalues (i.e. time-scales) and spatial symmetry of eigenfunction are consistent with symmetries of the potential energy, and eigenvalue and eignfunction errors are reported in the header.
Figure 9 : Third triplet of leading eigenfunctions of the Langevin generator for a cobalt nanoparticle potential. First row show 3D visualization on the unit sphere S2 , colored by eigenfunction value, while the second one shows polar coordinate representation ψ(θ,φ) . Third row shows reult obtained by LITL. Multiplicity of learned eigenvalues (i.e. time-scales) and spatial symmetry of eigenfunction are consistent with symmetries of the potential energy, and eigenvalue and eignfunction errors are reported in the header.
Figure 10 : Accuracy–fairness Pareto curve for the Adult Income experiment. Each point corresponds to a different number of LITL steps. LITL hyperparameters are calibrated on audit set, and then used to achieve substantial reduction in TPR disparity with a limited decrease in accuracy on a test set.
Figure 11 : Training dynamics of the LITL encoder. Physics-informed contrastive loss measurign the discrepancy of
Figure 12 : Accuracy deterioration and TPR disparity improvement across LITL gradient steps on 100 random test datasets. On the left the LITL model uses 16 slowest directions, while on the right the flow is across additional 16 faster ones. While slowest manifold corrects the TPR gap while minimally changing the accuracy it reaches the metastable state at around 50% gap reduction. On the other hand, including faster dynamics allows further reduction, but the transformations of the latent space geometry induce accuracy deterioration. The gap in the spectrum shown in the middle plot shows how the two manifolds are separated. This demonstrates how spectral structure of the generator related to the feedback potential allows one to introduce alignment of the model in a control manner.
Molecular dynamics simulations proceed by integrating the Langevin equations over many small femtosecond timesteps. This poses a challenge for estimating ensemble properties and transition dynamics that occur on much longer timescales. We introduce Langevin Flow Maps, which extend machine-learned force-fields to additionally learn the stochastic Langevin integrator. We show that Langevin Flow Maps enable large-timestep molecular dynamics and recover accurate dynamical properties of the system, while running an order of magnitude faster than current machine-learned force fields. Further, by training on a diverse molecular dataset, we demonstrate a path towards transferable Langevin Flow Maps.
We propose a spectral learning method for stochastic nonlinear dynamical systems represented with embedded latent transfer operators in deep feature spaces. We instantiate the method as Deep Spectral Encoder (DSE), an operator-based latent state-space model in which a time-invariant neural encoder implements learnable nonlinear feature maps from observations, and these features define Markovian latent states whose temporal evolution and observation mapping are described by the transfer and observation operators, respectively. Functional canonical correlation analysis in a learnable Galerkin-projected feature space provides state coordinates from past and future observations, and the two linear operators are estimated on the state coordinates as ridge-regularized closed-form solutions that coincide with Galerkin projections of the associated covariance operators. On this representation, we generalize sequential Bayesian filtering and Koopman spectral mode decomposition in feature space. Experiments on several scenarios show stable and superior performance with sequential Bayesian filtering and dynamic mode decomposition baselines even under noise and partial observability.
Ryogo Tanaka, Yoshinobu Kawahara
Graduate School of Information Science and Technology, The University of Osaka, Osaka, Japan · Center for Advanced Intelligence Project, RIKEN, Tokyo, Japan
We study Slowly Annealed Langevin Dynamics (SALD), a sampler for tracking a path of moving target distributions and approximating the terminal target through time slowdown. We establish non-asymptotic convergence guarantees via a KL differential inequality, showing that slowdown improves tracking through contraction of intermediate targets and the complexity of the path. Motivated by training-free guided generation with pretrained score-based generative models, we further introduce Velocity-Aware SALD (VA-SALD), which explicitly incorporates the underlying marginal distributions of the pretrained model and uses slowdown to correct the additional deviation induced by guidance. This yields a principled framework for training-free guided generation for diffusion-based and related generative model families, together with convergence guarantees that clarify the roles of intermediate functional inequalities and guidance bias. Code is available at https://github.com/anitan0925/sald.
Atsushi Nitanda, Dake Bu, Yueming Lyu +1
Agency for Science, Technology and Research (A⋆STAR) · Nanyang Technological University · City University of Hong Kong