Gaussian processes (GPs) are widely used across machine learning, spatial statistics, time-series analysis, optimization, Bayesian statistics, and scientific applications. A central component of a GP model is its kernel, which is typically specified through a parametric family. Among the most widely used choices is the radial basis function (RBF), also known as the squared exponential or Gaussian kernel, owing to its simple form, smoothness, and flexibility. In practice, the kernel parameters are routinely estimated by the maximum likelihood estimators (MLEs), as implemented by standard GP software. Despite this widespread use, the asymptotic behavior of the MLEs remains poorly understood under fixed-domain asymptotics, even for the RBF kernel. The main difficulty arises from the increasingly strong dependence among densely sampled observations and the nonlinear dependence of the covariance matrix on the kernel parameters. In this paper, we address this gap by providing, to the best of our knowledge, the first complete asymptotic characterization of the joint MLE of the spatial variance, lengthscale, and nugget variance under fixed-domain asymptotics. We establish consistency, derive convergence rates for all three parameters, prove joint asymptotic normality, and show that these rates are minimax optimal.
Figures & tables
Figure 1: Simulation results for p=1 on the log scale, with θ0=(1,0.25,0.01) and 1000 replicates per n . Top: RMSE (filled circles), robust standard deviation IQR/1.349 (open squares), both with bootstrap 95% intervals, and the exact asymptotic standard deviation [In(θ0)−1]jj1/2/θ0j (grey line). The grey region indicates sample sizes beyond the Monte Carlo range, where repeated MLE computation is computationally infeasible but the Fisher-information can still be evaluated. Dashed lines show the theoretical rates in Theorem 3.1 (ii). Bottom: normal Q–Q plots of the standardized errors, with a 95% band for exact normality; triangles mark points outside the frame.
Figure 2: Simulation results for p=2 , with the layout and graphical elements as in Figure 1 .
Figure 3: Simulation results for p=3 , with the layout and graphical elements as in Figure 1 .
RMSE
IQR/1.349
Asymptotic SD
Theory
logσ2
p=1
−0.49
−0.49
−0.35
−0.5
p=2
−0.75
−0.72
−0.74
−1
p=3
−1.34
−1.48
−1.32
−1.5
logl
p=1
−2.05
−1.81
−1.66
−1.5
p=2
−2.32
−2.24
−2.22
−2
p=3
−3.11
−3.11
−3.09
−2.5
Table 1: Least-squares slopes of log(error) against logbn for logσ2 and logl , and against logn for logτ2 , over the nine simulated sample sizes ( 102≲n≤106 ). The last column gives the theoretical rate shown in Theorem 3.1 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Simulation results for p=1 on the original parameter scale, with the layout and graphical elements as in Figure 1 .
Figure A2: Simulation results for p=2 on the original parameter scale, with the layout and graphical elements as in Figure 1 .
Figure A3: Simulation results for p=3 on the original parameter scale, with the layout and graphical elements as in Figure 1 .
We study ridge-regularized log-density-ratio estimation in the Gaussian location model with a common covariance matrix. By affine invariance, the model is written as q ∼ N(0, I), p ∼ N(Δ, I), with linear features, where Δ is a mean vector. The variational estimator is the empirical Kullback-Leibler (KL) log-normalized fit with a squared L2-penalty on its nonconstant coefficient, and the spectral estimator recently introduced in [1] replaces a single variational problem by a continuum of ridge-regularized least-squares problems. We derive high-dimensional deterministic asymptotic equivalents when the numbers of observations and dimension tend to infinity with fixed ratios. The regularized variational limit is characterized by a scalar entropy minimization problem derived from the convex-Gaussian-min-max theorem (CGMT), while the regularized spectral limit follows from deterministic equivalents for resolvents of weighted sums of two independent Gaussian sample covariance matrices. We use these formulas to compare population risks, with experiments focused on fixed-signal aspect-ratio sweeps and optimized regularization. Our conclusion is that with many observations, under the criteria and asymptotic regimes analyzed here, the well-specified variational estimator has the smaller risk, while with fewer observations, the spectral estimator is favored because its covariance-based construction has lower variance. We also study how a nuclear penalty can be used and partially analyzed to perform feature learning.
Francis Bach
SIERRA · Inria - Ecole Normale Sup´erieure PSL Research University
We show that, up to isotropic scaling, the Gaussian RBF reproducing kernel Hilbert space (RKHS) is asymptotically isometric to Euclidean space in the large bandwidth limit. This strongly suggests that kernel-based constructions reliant on metric properties of the RKHS will yield results for Gaussian RBF kernels that similarly approach those of linear kernels for large bandwidths. The asymptotic behavior of Gaussian CKA can be understood in this light. We further consider kernel PCA, showing that Gaussian RBF eigenvalues, eigenprojections, and principal components all converge to those of classical (linear) PCA as bandwidth σ→∞. For a given data representation, both the RKHS feature embeddings and the orthogonal PCA eigenframes of the two kernel types differ asymptotically by a geometric similarity transformation, up to a residual of size O(σρ)2, where ρ is a measure of geometric eccentricity of the representation, equal to the ratio of maximum to median pairwise distance between data examples. Experiments over a diverse collection of data sets demonstrate that ρ provides a simple and reliable predictor of dataset-specific convergence behavior in the top principal directions.
Sergio A. Alvarez
Department of Computer Science, Boston College, Chestnut Hill, MA 02467 USA
The nonparametric maximum likelihood estimator (NPMLE) of a Gaussian location mixture maximizes the likelihood over the infinite-dimensional space of mixing distributions. The maximizing mixing distribution can be nonunique, and the classical bound on its number of atoms grows linearly with the sample size n. We show that a vanishingly small random perturbation of the likelihood yields exact polylogarithmic sparsity. The resulting randomly reweighted NPMLE maximizes a weighted likelihood whose independent weights, taken to be Gamma in our analysis, concentrate around one as n grows. With high probability, it is unique, has O{(logn/loglogn)d+logn} atoms in dimension d, nearly maximizes the ordinary likelihood, and estimates the mixture density at a Hellinger rate that is parametric up to logarithmic factors. This sparsity holds for the estimator itself, not for an approximation of it, and requires no support penalty. The proof rests on an effective-dimension principle for positive kernel mixtures: low-dimensional variation of the fitted values controls the support of every extreme point of the set of maximizers. Numerical illustrations verify that the reweighted NPMLE has Hellinger risk and support size comparable to those of the ordinary NPMLE.