A Polyphonic Conception of AI Understanding
Organizations: University of Bern, Department of Philosophy · École Polytechnique Fédérale de Lausanne (EPFL) · Idiap Research Institute · Machine Alignment, Transparency, and Security (MATS)
Abstract
When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but name. Yet the question is ill-framed as it stands, because the inherited concept operates within a monophonic paradigm: the idea that a cognitive system's understanding of something must be localised to a single mechanism underpinning all the capacities conferred by such understanding. Drawing on a wide range of mechanistic evidence, we show that LLMs are pervasively polyphonic: outputs emerge from coalitions of parallel mechanisms of uneven reliability, which variously complement, duplicate, or drown out one another, with several coalitions sufficing for a task without any one being indispensable. Polyphony not only complicates attributions of understanding, but renders monophonic inference patterns hazardous. In response, we develop a conception of understanding fit for polyphonic AI. It centres on sound circuitry that is reliably and correctly recruited and in control of outputs. Attributions of understanding thereby become tractable claims about internal organisation, and can do the work of guiding trust in AI.
Figures & tables
| Jury room element | LLM component |
|---|---|
| Communal table | Residual stream |
| A note on the table | A feature |
| Each day and panel of jurors | One layer |
| New jury each day | No memory between layers (except what is in the residual stream) |
| Selecting evidence | Attention heads |
| Processing evidence, recalling information | MLPs |