cs.AISep 7, 2026

The Internal Anatomy of Strategic Choice in Large Language Models

Authors: Vinícius FerrazLeon HoufEnrico Ferrea

Abstract

Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal 2×22\times2 games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.

Explore similar work

CardsList
  1. Probing Persona-Dependent Preferences in Language Models

    May 13, 2026Oscar Gilg, Pierre Beckmann, Daniel Paleka +1Pairwise Preferences