stat.MLOct 8, 2026

Feature Space Adaptation for Effortless Gaussian Process Flows

Authors: Thomas Cowperthwaite, Louis Sharrock, Lachlan Astfalck, Henry Moss

Organizations: University of Cambridge · University College London · University of New South Wales · Lancaster University

Abstract

Outside the linear-Gaussian regime, conditional sampling from Gaussian processes (GPs) is challenging. Recent methods such as FlowGP (Moss et al., (2026)) can condition on arbitrary non-linear and non-Gaussian statements, but at considerable cost: an expensive iterative and high-dimensional diffusion that requires hand-specified kernel hyperparameters. In this paper, we alleviate two significant drawbacks of FlowGP by (1) introducing kernel approximations that enable scaling to high-resolution domains and (2) proposing a way to obtain the marginal likelihood by measuring the work needed to steer the diffusion towards conditioning statements. We enable, for the first time, hyperparameter optimisation within FlowGP and demonstrate our approach on probabilistic downscaling from areal summary statistics, PDE solution inference on irregular domains, and recovery of sea level anomaly fields from non-Gaussian satellite observations.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 20, 2026stat.ML

Conditioning Gaussian Processes on Almost Anything

Gaussian processes (GPs) offer a principled probabilistic model over functions, but exact inference is restricted to the linear-Gaussian regime. We establish an explicit equivalence between GPs and a class of linear diffusion models, recasting predictive sampling as an ODE with closed-form Gaussian dynamics and a likelihood-dependent guidance term that admits a simple Monte Carlo approximation. In the linear-Gaussian setting, we recover standard GP conditioning exactly; beyond conjugacy, the same machinery handles any conditioning statement admitting point-wise likelihood evaluation -- including non-linear physics, and, for the first time, natural language via large language models. Whitening isolates the irreducible non-Gaussian dynamics, minimising Wasserstein-2 transport cost and eliminating numerical stiffness. The result is a general-purpose GP inference scheme requiring no bespoke derivations. Together, these results provide a general mechanism for incorporating the full richness of real-world knowledge as conditioning information, opening a new frontier for the probabilistic modelling of real-world problems.
Sep 28, 2026cs.LG

What Does a Stream Model Buy You in Flow Matching?

Stream-level flow matching replaces the linear interpolant of conditional flow matching (CFM) by a Gaussian-process (GP) stream connecting each source--target pair, and reports lower sample error than \icfm{} on 2-Gaussian, MNIST and CIFAR-10 benchmarks. We ask what such a stream model actually contributes. Three results answer the question. (i)\emph{Reduction.} The stream-level CFM objective depends on the stream law only through the per-time joint law of (st,⋅t)(s_t,\sdot_t), so the conditional paths a Gaussian stream can reach are exactly the Gaussian conditional paths CFM already parametrises; in the coordinate-wise, shared-scalar-kernel construction gpcfm actually uses, the entire design space collapses to two scalar curves (mt,vt)(m_t,v_t), and cross-time covariance affects only estimator variance. (ii)\emph{The GP is a constrained chart of that space.} One kernel sets both mtm_t and vtv_t, so the paper's own recipe for widening coverage-shrinking the SE length-scale---destroys the interpolant (the midpoint mean weight falls from 1.031.03 to 0.000.00). On the 2-Gaussian benchmark this makes the GP chart diverge on 15/20015/200 runs at high coverage against 0/2000/200 for a decoupled (mt,vt)(m_t,v_t) chart (p=6.6×10−5p=6.6\times10^{-5}), and crossing the two curves shows the divergence tracks the mean, not the variance. On MNIST the same sweep does not diverge and the ordering reverses, so whether the coupling is harmful is benchmark-dependent; what holds on both is that the recipe buys nothing---no coverage level beats the paper's own, and past max⁡tvt≈0.6\max_t\sqrt{v_t}\approx0.6 both charts degrade. (iii)~\emph{Audit.} The released code does not implement the mechanism it describes: state and velocity are drawn independently (corr=0.00±0.01\mathrm{corr}=0.00\pm0.01 against an intended ±0.83\pm0.83--0.990.99).
May 11, 2026stat.ML

Scalable Gaussian process inference via neural feature maps

We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels. We show that the learned feature map can be interpreted as an optimal low-rank approximation to a Gram matrix derived from an implied RKHS, from which we establish consistency of the GP posterior. We further analyse the spectral properties of the induced kernels and introduce product feature-map kernels to address oversmoothing. This simple yet powerful approach enables fast, scalable, and accurate exact GP inference with minimal upfront work. The flexibility of kernel design supports seamless application to both regression and classification tasks across diverse data modalities, including tabular inputs and structured domains such as images. On benchmark datasets, this approach surpasses pre-existing methods in terms of accuracy and training and prediction efficiency.