From cacophony to hierarchy: a principled framework for assessing AI consciousness
Organizations: Google DeepMind, DeepMind Institute · Flourishing Intelligence Program, Centre for Eudaimonia and Human Flourishing, Linacre College, University of Oxford · Institute of Philosophy, University of London · Fitzwilliam College, University of Cambridge · AI Cognition Institute · Rethink Priorities · School of Engineering and Informatics, University of Sussex · Department of Brain Sciences, Imperial College London · Sussex Centre for Consciousness Science, University of Sussex · Leverhulme Centre for the Future of Intelligence, University of Cambridge · School of Computing, Australian National University · University College London · Department of Computing, Imperial College London · Department of Psychiatry, University of Oxford · LIFE · Wellcome Centre for Human Neuroimaging, University College London · Interacting Minds Centre, Aarhus University
Abstract
The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a cacophony of competing theories that often talk past each other. Separating the hard problem from the mapping problem allows the deepest metaphysical disagreements to be set aside: granting that experience supervenes on a system's organisation, the tractable question becomes at which grain of description that supervenience base sits. We extend Marr's three levels of analysis into a five-level hierarchy of functional descriptions (behavioural, computational, intrinsic causal-structural, organismic, and organism-environment) grounded in supervenience, coarse-graining, and multiple realisability. The major theories of consciousness are positioned within this hierarchy according to which level they take to be critical, and for each level we develop operationalisable indicators and assess current AI systems against them. A Bayesian model then combines theoretical credences with indicator evidence into an overall credence in a system's capacity for consciousness. In illustrative assessments, the verdict for current LLMs is driven as much by where theoretical credence is placed as by how the evidence is read: under different stipulated readings and credence distributions, assessments range from below 0.01 to roughly 0.8, showing sensitivity to assumptions. Finally, the consciousness indicators at each level closely overlap with the architectural features needed for general intelligence, suggesting that increasingly capable AI may become a stronger candidate for consciousness. The framework supports a structured agnosticism, in which theoretical commitments are made explicit, credences are updated as evidence accumulates, and assessments take the form of aggregated probabilities rather than verdicts.
Figures & tables
| MARR'S LEVEL | OUR LEVEL | DESCRIPTION |
|---|---|---|
| Computational | Level 1: Behavioural | Input-output function; what the system does |
| Algorithmic | Level 2: Computational functional | Algorithms and information processing; how the system does it |
| Implementational | Levels 3, 4, 5 | Intrinsic causal structure, organismic functions, organism-environment coupling |
| INDICATOR | DESCRIPTION | THEORY LINK |
|---|---|---|
| Information integration | Widespread information sharing; high mutual information between subpartitions; the system becomes a globally interdependent object both cross-sectionally and temporally. The whole must be informationally 'more than the sum of its parts'. | GWT, CF-IIT, PPT, HOTT, RPT |
| Recursivity | Widespread recursive informational flows; feedback reentrant loops; algorithmic recurrence | RPT, PPT, GWT, CF-IIT |
| World model | An internal causal representation of the external environment with a smooth representational space and meta-representations. | PPT, HOTT |
| Self model | A representation of being a distinct unified entity that owns its parts, directs attention, and persists through time, which is created by the system to track and predict the world. | PPT, link to Level 5 |
| Attentional competition and meta-attention | Informational contents, 'ideas', representations etc are amplified or attenuated through recurrent processing with higher order relevance influencing the amplification; effectively forming a routing bottleneck for global availability of information. Meta-attention and attention modelling reinfluences attention. | RPT, GWT, PPT, AST |
| Meta-cognition | Higher-order processes that monitor, model and understand, and control lower level processes; e.g. meta-cognitive knowledge (including meta-representations) and meta-cognitive regulation (executive control). Introspectivity. | HOTT, AST, PPT, GWT |
| INDICATOR | LEVEL 2 FORMULATION | LEVEL 3 REFORMULATION |
|---|---|---|
| Information integration | Algorithmic mutual information between subsystems | Genuine physical causal integration: physical components constrain each other, creating a causally integrated whole |
| Recursivity | Algorithmic feedback loops | Genuine physical causal reentrance: simultaneous reciprocal causal influence between physical components, not merely sequential feedback |
| World model | Computed internal representation of causal structure | Causally constitutive world model: the system's physical organisation itself constitutes a model of the environment, not merely stores one |
| Self model | Computed representation of the system's own states and capabilities | Causally embedded self-model: physical components whose causal role is to track and constrain the causal dynamics of the rest of the system |
| Attentional competition and meta-attention | Algorithmic selection mechanism that amplifies and suppresses information | Genuine causal competition: different physical components compete for influence over the system's global state through intrinsic causal dynamics, with outcome emerging from interaction not from a selection algorithm |
| Metacognition | Monitoring algorithm that evaluates the system's own processing | Causally integrated metacognition: monitoring and processing are simultaneous, reciprocally constraining aspects of the same causal dynamics, not sequential sampling and evaluation |
| INDICATOR | DESCRIPTION | IMPLICATION FOR AI |
|---|---|---|
| 1. Existential Precariousness | The system would need to have a continued existence that is not guaranteed, that requires active work to maintain. Its processing would need to be oriented toward maintaining its own viability, and this orientation would need to be an intrinsic feature of the system's organisation rather than a programmed goal; the system maintains itself because its architecture is such that self-maintenance is what the system does. | Current AI systems completely lack this property. A transformer has no stake in its own continued operation. It does not need to actively maintain itself. Its computations are not oriented toward self-maintenance because there is nothing to maintain; the system's existence is guaranteed by external infrastructure that the system itself has no relationship to. This may change in future embodied systems like radically self-sufficient robots. |
| 2. Homeostatic or allostatic regulation | The system would need to continuously monitor and adjust its own internal parameters to keep them within viable ranges, and these regulatory processes would need to be causally integrated with the system's information processing. | This goes far beyond current approaches to AI self-monitoring. It would require something more like a system that has 'metabolic' needs, e.g. that actively manages its energy supply, its thermal state, its hardware integrity, and its memory allocation, in a way that is not handled by external infrastructure but is the system's own ongoing concern, integrated into its cognitive processing. |
| 3. Interoceptive inference | The system would need to generate internal perceptions of its own regulatory state, and these internal perceptions would need to have the character of predictions or inferences about the system's own condition. | For an AI system, this would require having a predictive model of the system's own internal dynamics that generates expectations about what the system's internal state should be and that registers surprise when it deviates from expectation. |
| 4. Valenced affect | The system's interoceptive inferences would need to have an inherent motivational character, a felt quality of better or worse that is grounded in the system's relationship to its own viability. | The system would need to have states that are intrinsically better or worse for its continued viable operation, and its interoceptive inference system would need to track these states in a way that generates motivational force — not because a reward function says so but because the system's own architecture makes self-maintenance intrinsically rewarding and self-destruction intrinsically aversive. |
| 5. Embodied agency | The system would need to act in the world in ways that are driven by its interoceptive states and oriented toward maintaining its viability. It would need to be not merely an information processor but an agent whose actions are motivated by its felt relationship to its own condition. | The link between interoceptive inference, valenced affect, and action would need to be tight and constitutive, not mediated by an external action-selection algorithm but emerging from the system's own regulatory dynamics. |
| 6. Physical state-sensing | The system would need to have components whose physical state changes in real time in response to changes in the system's overall condition, in a way that is causally efficacious for the system's subsequent processing (cf. ionic gradients across cell membranes). | The physical organisation of the system would need to change in response to the system's state, and this change would need to constitute the system's awareness of its own condition. This is not information processing in the usual sense; rather it is physical self-transformation as a mode of self-sensing. |
| INDICATOR | DESCRIPTION | IMPLICATION FOR AI |
|---|---|---|
| 1. Sensorimotor coupling | The system would need to be embedded in an environment through sensory and motor channels that are continuously and bidirectionally linked, such that the system's actions systematically change its sensory input and its sensory input systematically guides its actions. | Current AI systems almost entirely lack this property. A language model has no sensorimotor coupling at all. Robotic systems come closer, but even here the coupling is often impoverished. |
| 2. Mastery of sensorimotor contingencies | The system would need to have not merely been subject to structured sensorimotor contingencies but to have learned the structure of those contingencies and to use this knowledge in its ongoing interaction with the environment ( Noë, 2004 ) . | This would require evidence that the AI system has internalised the structure of its sensorimotor contingencies: that it can anticipate how its actions will change its sensory input, that it is surprised when the contingencies are violated, and that its perceptual discriminations are grounded in practical knowledge of how things behave under interaction. |
| 3. Affordance responsiveness | The system would need to perceive its environment not merely as a collection of objects with properties but as a field of possibilities for action — affordances. These affordances are relational properties that depend on both the environment and the organism's embodiment. | Affordance-responsiveness would mean that the AI system perceives its environment in terms of action possibilities specific to its own embodiment. The perceptual world shifts in response to the system's own condition and capacities. |
| 4. Stable embodied perspective | The system would need to have a coherent, continuous viewpoint on the world that is anchored in its body and that provides a stable spatial and temporal framework for its experience – a here, a now, a facing-this-way, from which the world is organised. | This would require that the AI system's sensory and motor coupling with the environment generates a coherent egocentric frame, continuously updated by the system's movements and actions. |
| 5. Environmental embedding and constraint | The system would need to be subject to the constraints of its environment in a way that shapes its processing and behaviour — subject to physical forces, to the passage of time, to the limitations of its own body, to the resistance and unpredictability of the physical world. | The system's processing would need to be shaped by its real-time engagement with a physical environment that imposes delays, noise, unpredictability, and resistance. A simulation may not suffice because the constraints are programmed parameters that mimic constraint without having the same direct causal force. |
| 6. Enactive autonomy | The system would need to be not merely responsive to its environment but actively generating and maintaining its own patterns of sensorimotor engagement. In biological organisms, this is deeply tied to autopoiesis — the self-production of the organism's own organisation. | Genuine autopoiesis might require something like an artificial organism. But a lower bar would entail that the system itself should have some capacity to shape its own sensorimotor engagement, to develop new sensory strategies, to discover new action possibilities. |
| LEVEL | THEORIES OPERATING AT THIS LEVEL |
|---|---|
| Level 1: Behavioural | Analytic behaviourism |
| Level 2: Computational functional | GWT, HOTT, RPT, AST, PPT, CF-IIT |
| Level 3: Intrinsic causal-structure functional | IIT, ICS-RPT, ICS-GWT, ICS-PPT |
| Level 4: Organismic functional | Biological Naturalism (Searle), Beast Machine Theory (Seth), Biopsychism |
| Level 5: Organism-environment (4E) functional | Enactivism, Sensorimotor contingency theory (Noë), Ecological psychology (Gibson), Extended mind thesis (Clark, Chalmers) |
| Substrate-dependent realisability constraints (cross-cutting) | EM field theories (McFadden) Level 3; Quantum theories (Penrose-Hameroff, Neven) Level 3; Carbon chauvinism Level 4; Biological Naturalism (substrate dimension; Searle/Seth) Levels 3 & 4; Block's meat hypothesis Levels 2, 3 & 4; Lane's ionic gradients Level 4; Large-scale dynamic patterns (Godfrey-Smith) Level 3 |