cs.AISep 20, 2026

Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer

Authors: David Litman

Organizations: Computational Neurobiology Laboratory, Salk Institute

Abstract

Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer takes Elo, a measure of competitive chess skill, as an input to the pre-trained network, so we can vary the skill the network is conditioned on with no change to its weights. Here we investigate how turning this skill dial affects self-attention. Ablating every attention head at every Elo from 700 to 2500, we find 1) increasing skill pushes the causal center of mass of the computation deeper, monotonically, for every chess piece and move type we measured; 2) the depth migration is much greater for specific tactics, especially knight forks, than for other move types; 3) the migration consists of deeper heads getting recruited for more specialized computations while one shared shallow head keeps a roughly constant contribution. These results may shed light on how conditioning inputs redistribute computation in larger transformers.

Figures & tables

Explore similar work

CardsList
  1. Chessformer: A Unified Architecture for Chess Modeling

    May 18, 2026Daniel Monroe, George Eilender, Philip Chalmers +2ChessTransformer Encoder

  2. Depth-Attention: Cross-Layer Value Mixing for Language Models

    Jun 3, 2026Boyi Zeng, Yiqin Hao, Zitong Wang +7Layer-WiseCross-Layer Interactions

  3. Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

    May 28, 2026Benjamin Walker, Terry LyonsChessLatent States