cs.LGOct 7, 2026

Leaner Transformers Can Easily Learn to Cluster

Authors: Charlotte Park, Kenneth L. Clarkson, Lior Horesh, Takuya Ito, Parikshit Ram

Organizations: MIT · IBM Research

Abstract

Transformers have in-context learning capabilities, where some known learning algorithms can be executed in the forward pass through the model. Recent work shows that transformers can exactly perform Lloyd's algorithm for kk-means clustering with nn points in dd dimensions with an embedding size demb=d+kd_{\textsf{emb}} = d+k (thus, requiring attention projection matrices of size (d+k)2(d+k)^2). In this work, we build upon this result in the following ways: First, we present an equally expressive but smaller transformer that executes Lloyd's algorithm with embedding size demb=(d+⌈log⁡2k⌉)d_{\textsf{emb}} = (d + \lceil \log_2 k \rceil). Next, we train these transformers to learn the clustering algorithms given a distribution of clustering tasks, and theoretically characterize and empirically validate the factors affecting the convergence and in-distribution generalization of learning algorithms based on stochastic gradients. Finally, we probe the general clustering abilities of these learned algorithms (in the form of transformers), and try to understand situations where they succeed and fail.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent

    May 7, 2026Chenyang Zhang, Yuan CaoIn-Context LearningTransformer Architectures

  2. Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers

    May 8, 2026Noboru Isobe, Daisuke Inoue, Masaaki ImaizumiTransformer AttentionTransformer Architectures

  3. Fixed Universal Transformers

    May 29, 2026Jingwen Liu, Alexandr Andoni, Daniel HsuTransformer ArchitecturesEnergy-Based Transformer