cs.LGOct 6, 2026

Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

Authors: Martin Eppert, Krishna Balasubramanian, Subhro Ghosh, Jason Klusowski, Yan Shuo Tan

Organizations: National University of Singapore · University of California, Davis · Princeton University

Abstract

Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from. We study how such adaptivity is learned in a controlled location-estimation problem. Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order n−1n^{-1}, whereas uniform data are best estimated from their extremes, at the faster rate n−2n^{-2}. We also provide the example of a symmetric Gaussian mixture, for which a rate of σn2/nσ^2_n/n can be attained. On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function. A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range. We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow. With Ω~(n1+ε)\widetildeΩ(n^{1+ε}) pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor nεn^ε of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime. These guarantees extend to new locations and longer contexts. A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast. Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection. End-to-end experiments recover the predicted specialization.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Context-Adaptive Inference: A Unified Statistical and Foundation-Model View

    Jul 25, 2026Yue Yao, Caleb N. Ellington, Jingyun Jia +9Parameter-Efficient AdaptationIn-Context Learning

  2. Adaptive Kernel Density Estimation with Pre-training

    May 13, 2026Ruitong Zhang, Ke DengDensity Ratio EstimationKernel Method

  3. Adaptive Bayesian Online Learning via Expert Aggregation

    Jul 22, 2026Jungbin Jun, Ilsang OhnLearning-Augmented AlgorithmsBayesian