cs.LGSep 30, 2026

Group-Invariant Statistics Determine Embedding Geometry: Harmonic Analysis of Representations from Bach to the Night Sky

Authors: Liam Storan, Andreas Tolias, Nina Miolane

Organizations: Stanford University · University of California, Santa Barbara

Abstract

The representations that language models learn for concepts such as months, weekdays, and places display consistent geometric structure: circles and saddle-shaped "Pringle" manifolds. Recent work traced these structures to translation symmetry\textit{translation symmetry} in word co-occurrence statistics, deriving the observed Fourier geometry when co-occurrence depends only on distance on an abelian lattice of concepts. We demonstrate that more general notions of symmetry lead to equally structured predictions. Considering symmetries defined by arbitrary finite groups, compact groups, and homogeneous spaces, we prove that whenever the co-occurrence statistics of a word family are invariant under a group GG, the learned word embeddings consist of matrix elements of the irreducible representations (irreps) of GG. Circles and Pringles arise when GG is cyclic, in which case the irreps are Fourier modes. We verify the irrep structure in three experimental settings. (i) The cyclic group Z12\mathbb{Z}_{12}: for the months of the year we recover the known circular geometry. (ii) A dihedral group acting on the major and minor triads: we unify two classical observations -- that transposition and chord inversion form a group (T/IT/I) acting on chords (music theory), which implies\textit{implies} that the well-known "circle of fifths" emerges in learned chord embeddings (machine learning). (iii) We explain and reproduce a recently discovered spherical representation of celestial objects in large language models (LLMs) as a spherical-harmonic embedding derived from our theory. Our results demonstrate that the geometry of learned representations is often a consequence of the statistical symmetry of underlying data.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Symmetry in language statistics shapes the geometry of model representations

    Feb 16, 2026Dhruva Karkada, Daniel J. Korchinski, Andres Nava +2SymmetryManifolds

  2. Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization

    May 12, 2026Zhehang Du, Hangfeng He, Weijie SuSymmetry

  3. Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence

    May 22, 2026Andres Nava, Matthieu WyartContextual EmbeddingsToken Co-Occurrence Graphs