cs.LGMay 7, 2026

Attributions All the Way Down? The Metagame of Interpretability

Authors: Hubert BanieckiPrzemyslaw BiecekFabian Fumagalli

Abstract

We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution φ(f)φ(f) explaining a model ff, we measure the directional influence of feature jj on the attribution of feature ii, denoted as meta-attribution φji(f)\varphi_{j \to i}(f), by treating the attribution method itself as a cooperative game and computing its Shapley value. Theoretically, we prove that attributions hierarchically decompose into meta-attributions, and establish these as directional extensions of existing interaction indices. Empirically, we demonstrate that the metagame delivers insights across diverse interpretability applications: (i) quantifying token interactions in instruction-tuned language models, (ii) explaining cross-modal similarity in vision-language encoders, and (iii) interpreting text-to-image concepts in multimodal diffusion transformers.

Explore similar work

CardsList
  1. MorphoSHAP: Rethinking the Unit of Attribution in Explanation for Deep Visual Models

    Sep 22, 2026Anirudh Prabhakaran, Alexandre Rocchi, Gianni Franchi