stat.MLMay 14, 2026
SaveLarge Dimensional Kernel Ridge Regression: Extending to Product Kernels
Abstract
Recent studies have reported and in large dimensional kernel ridge regression (KRR). However, these findings are predominantly derived under restrictive settings, such as inner product kernels on sphere or strong eigenfunction assumptions like hypercontractivity. Whether such behaviors hold for other kernels remains an open question. In this paper, we establish a broad, new family of large dimensional kernels and derive the corresponding convergence rates of the generalization error. As a result, we recover key phenomena previously associated with inner product kernels on sphere, including: the when the source condition ; the when ; a in the convergence rate and a with respect to the sample size .
Explore similar work
Conditionally positive definite (CPD) kernels are defined with respect to a function class . It is well known that such a kernel is associated with its native space (defined analogously to an RKHS), which in turn gives rise to a learning method -- called conditional kernel ridge regression (conditional KRR) due to its analogy with KRR -- where the estimated regression function is penalized by the square of its native space norm. This method is of interest because it can be viewed as classical linear regression, with features specified by , followed by the application of standard KRR to the residual (unexplained) component of the target variable. Methods of this type have recently attracted increasing attention. We study the statistical properties of this method by reducing its behavior to that of KRR with another fixed kernel, called the residual kernel. Our main theoretical result shows that such a reduction is indeed possible, at the cost of an additional term in the expected test risk, bounded by , where is the sample size and the hidden constant depends on the class and the input distribution. This reduction enables us to analyze conditional KRR in the case where is positive definite and is given by the first principal eigenfunctions in the Mercer decomposition of . We also consider the setting where consists of random features from a random feature representation of . It turns out that these two settings are closely related. Both our theoretical analysis and experiments confirm that conditional KRR outperforms standard KRR in these cases whenever the -component of the regression function is more pronounced than the residual part.
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction. Specifically, we consider the problem of recovering a multi-index polynomial , with and , from finitely many data/label pairs. Importantly, the target function depends on input only through the projection onto an unknown -dimensional central subspace. The algorithm we analyze is appealingly simple: fit kernel ridge regression (KRR) to the data and compute the Average Gradient Outer Product (AGOP) from the fitted predictor. Our main results show that under reasonable assumptions the top -dimensional eigenspace of AGOP provably recovers the central subspace, even in regimes when the prediction error remains large. Specifically, if the target function has degree , it is known that samples are necessary for KRR to achieve accurate prediction. In contrast, we show that if a low degree component of already carries all relevant directions for prediction, subspace recovery occurs in the much lower sample regime for any . Our results thus demonstrate a separation between prediction and representation, and provide an explanation for why iterative kernel methods such as Recursive Feature Machines (RFM) can be sample-efficient in practice.
Learning Curves and Benign Overfitting of Spectral Algorithms in Large Dimensions
Existing large-dimensional theory for spectral algorithms resolves either the optimally tuned point or the interpolation limit, but leaves the under-regularized regime unexplored. We study the learning curve and benign overfitting of spectral algorithms in the large-dimensional setting where the sample size and dimension are of comparable order, i.e., for some . We first consider inner-product kernels on the sphere and establish a sharp asymptotic characterization of the excess risk across the full regularization path under various source conditions , where measures the relative smoothness of the regression function. Our results reveal that the learning curve is not simply U-shaped but instead consists of three distinct regimes: over-regularized, under-regularized, and interpolation regimes. This characterization allows us to fully capture the benign overfitting phenomenon, demonstrating that benign overfitting arises consistently across both the under-regularized and interpolation regimes whenever is positive but no larger than a critical threshold. We further show that, in the sufficiently regularized regime, the kernel learning curve is recovered by an associated sequence model. Finally, we extend the learning-curve analysis to large-dimensional KRR for a class of kernels on general domains in whose low-degree eigenspaces satisfy spectral-scaling and hyper-contractivity conditions.