cs.LGNov 13, 2025

Fast Generalized Neural Tangent Kernel Statistics via Trace Estimation

Authors: James Hazelden, Balaaji Reddy Nagireddy, Eric Shea-Brown

Organizations: Applied Mathematics, University of Washington & The Allen Institute, Seattle, WA.

Abstract

The empirical state-space Neural Tangent Kernel (NTK) describes the local learning geometry of a finite-width neural network, but computing it explicitly is almost always impractical in terms of computation and memory costs. Here, we show that many useful NTK statistics that characterize, for example, the dimensionality of learned updates or how two models or learning rules relate, can instead be efficiently approximated to very high accuracy via matrix-free products using randomized trace estimation. Namely, we use Hutch++ to estimate the NTK trace, Frobenius norm, effective rank, and alignment. Furthermore, we show that the positive-semidefinite structure of the NTK yields one-sided estimators that require only forward- or reverse-mode automatic differentiation. We validate these estimators across MLPs, recurrent GRUs, and a natural-language Transformer with up to 410 million parameters, in which the state-space contains high-dimensional four-tensors. We demonstrate orders-of-magnitude speedups, with the fastest estimator in a given application depending on the ratio of parameter and state dimensions. Equipped with these estimators, we examine rich and lazy RNN training using hidden-state NTK alignment and use NTK alignment as a regularizer for data-scarce knowledge distillation. We find that this regularization can modestly improve generalization, especially in very data-scarce settings. Together, these results suggest state-space NTK diagnostics are practical even at large scales.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning

    May 9, 2026James Hazelden, Laura Driscoll, Eli Shlizerman +1Neural Tangent KernelGradient Descent

  2. Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime

    May 24, 2026Ruchirinkil Marreddy, Chaoyue LiuNeural Tangent KernelOverparameterization

  3. A Function-Space Dichotomy for Compositional Learning: Exponential Sub-Optimality of the Neural Tangent Kernel

    Jul 7, 2026Arkaprabha Ganguli, Emil ConstantinescuNeural Tangent KernelRectified Linear Unit Networks