cs.LGOct 7, 2026

Fully Interpretable Minimal Transformers: From Geometry to Algorithm

Authors: Raneem Mahajne, Toviah Moldwin

Organizations: Edmond and Lily Safra Center for Brain Sciences, The Hebrew University of Jerusalem

Abstract

We present a framework for building and interpreting minimal transformer models. By constraining a transformer's embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations. Embeddings, query/key/value transforms, attention outputs, residual streams, and decision boundaries can all be seen directly. Our central claim is that the learned geometry implies an algorithm; the arrangement of points and boundaries in R^2 can be read as a step-by-step procedure. We train a transformer on a simple task where it must produce the most recently observed even number whenever the '+' operator appears in a sequence of digits. Once trained, we visually walk through every step of the transformer's computation. We show how the model embeds the tokens and their respective positions in the sequence, transforms them via the Q, K, and V matrices, uses the dot product between the Q and K representations to form the attention matrix, and uses the attention matrix to select values that move the representation of each input token to the region of the domain of the output layer that will correctly predict the next token. We introduce a suite of interpretability visualizations that make the algorithmic interpretation of this procedure explicit. Our framework offers a pedagogical and experimental testbed to explore how transformers use informational geometry to implement next-token prediction.

Figures & tables

Explore similar work

CardsList
  1. Transformers Linearly Represent Highly Structured World Models

    May 13, 2026Roman Kniazev, Nathanaël FijalkowTransformer ArchitecturesCombinations

  2. Trajectory Geometry of Transformer Representations Across Layers

    Jun 8, 2026Vishal Pandey, Gopal Singh, Yacine MahdidTransformer ArchitecturesLayer-Wise

  3. Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task

    Oct 21, 2025Brady Bhalla, Honglu Fan, Nancy Chen +1Transformer ArchitecturesBehavioral Embeddings