cs.CLSep 28, 2026

Causal and Interpretable Structures in LLM Compositional Tasks

Authors: Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Raphaël Sarfati, Christopher J. Earls

Organizations: Cornell University, USA · Goodfire AI, USA

Abstract

Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. For What Reason? Interpreting Models' Encoding of Causation and Antithesis

    Jul 20, 2026Abhidip Bhattacharyya, Shira WeinTransformer ArchitecturesCausal

  2. Large Language Models Develop Belief State Geometry In-Context

    Sep 15, 2026Daniel Balcells, Andrew Jun Lee, Chirag Rastogi +3In-Context LearningIn-Context