Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning
Organizations: School of Mathematics, Georgia Institute of Technology, Atlanta, GA 30332, United States · Department of Mathematics, Purdue University, West Lafayette, IN 47907, United States · Department of Mathematics and Halıcıoğlu Data Science Institute, University of California, San Diego, La Jolla, CA 92093, United States
Abstract
Transformers achieve remarkable performance by jointly learning broad families of tasks during pretraining and adapting to unseen tasks from only a short prompt. Yet a rigorous mathematical and statistical understanding of this phenomenon remains limited. This paper aims to study how Transformers exploit shared cross-task structure and how this structure affects the sample complexity of in-context learning (ICL). Specifically, we characterize task-space complexity through covering numbers under a prescribed metric, thereby quantifying the low-dimensional cross-task structure without requiring an explicit parametric representation. The resulting cover provides a set of anchor functions, which we use to introduce a task-identification-and-evaluation procedure: context observations localize an unseen task among the anchor functions, and the response at a query is predicted by aggregating the corresponding anchor function query evaluations. For approximation, we explicitly construct a Transformer with Softmax attention to approximate this procedure. For generalization, we derive an error bound that separates the effects of the number of pretraining tasks and the prompt length. The scaling with respect to the number of pretraining tasks is governed by the intrinsic dimensions of the task space and input domain; once sufficiently many tasks are available, the dependence on the prompt context length becomes dimension-free. To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL. Our theory provides a quantitative explanation of how joint pretraining across related tasks improves in-context generalization.
Figures & tables
| ICL on full family | ICL on 2D subspace | 21D ridge regression for 2D subspace | |
|---|---|---|---|
| 4 | 0.66663 | 0.07832 | 0.67697 |
| 8 | 0.50026 | 0.03092 | 0.47792 |
| 16 | 0.35547 | 0.01718 | 0.16517 |
| 20 | 0.26714 | 0.01623 | 0.03750 |
| Work | Task-space assumptions | Input and observation assumptions | Generalization error bound |
|---|---|---|---|
| This work | A uniformly bounded, uniformly -Hölder task space with , covering complexity . | Inputs sampled i.i.d. from an arbitrary distribution supported on a compact with ; noiseless responses. | |
| Kim et al. ( 27 ) | A Besov ball with ; under a B-spline wavelet expansion, the task coefficients are centered, independent, and satisfy a prescribed scale-dependent variance decay. | Inputs sampled i.i.d. from a distribution with density bounded above and below on ; bounded, mean-zero observation noise. | |
| Shen et al. ( 42 ) | Uniformly bounded -Hölder functions with , defined on a compact -dimensional Riemannian manifold with positive reach. | Inputs sampled i.i.d. from the uniform distribution on the manifold; noiseless responses. | |
| Ching et al. ( 10 ) | Tasks drawn from a distribution supported on a uniformly bounded -Hölder ball on , with a common . | Inputs sampled i.i.d. from a distribution with density bounded above and below on ; bounded observation noise satisfying . | |
| Hsu et al. ( 24 ) | A uniformly bounded function class within distance of a fixed finite-dimensional polynomial space with bounded coefficients. | Inputs sampled i.i.d. from a distribution supported on a bounded interval, with uniformly well-conditioned feature covariance; noiseless responses. |