From cacophony to hierarchy: a principled framework for assessing AI consciousness
Authors: Shamil Chandaria, Arvo Muñoz Morán, Fernando Rosas, Anil Seth, Henry Shevlin, Marcus Hutter, Thore Graepel, Adam Bales, +6 more
Organizations: Google DeepMind, DeepMind Institute · Flourishing Intelligence Program, Centre for Eudaimonia and Human Flourishing, Linacre College, University of Oxford · Institute of Philosophy, University of London · Fitzwilliam College, University of Cambridge · AI Cognition Institute · Rethink Priorities · School of Engineering and Informatics, University of Sussex · Department of Brain Sciences, Imperial College London · Sussex Centre for Consciousness Science, University of Sussex · Leverhulme Centre for the Future of Intelligence, University of Cambridge · School of Computing, Australian National University · University College London · Department of Computing, Imperial College London · Department of Psychiatry, University of Oxford · LIFE · Wellcome Centre for Human Neuroimaging, University College London · Interacting Minds Centre, Aarhus University
The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a cacophony of competing theories that often talk past each other. Separating the hard problem from the mapping problem allows the deepest metaphysical disagreements to be set aside: granting that experience supervenes on a system's organisation, the tractable question becomes at which grain of description that supervenience base sits. We extend Marr's three levels of analysis into a five-level hierarchy of functional descriptions (behavioural, computational, intrinsic causal-structural, organismic, and organism-environment) grounded in supervenience, coarse-graining, and multiple realisability. The major theories of consciousness are positioned within this hierarchy according to which level they take to be critical, and for each level we develop operationalisable indicators and assess current AI systems against them. A Bayesian model then combines theoretical credences with indicator evidence into an overall credence in a system's capacity for consciousness. In illustrative assessments, the verdict for current LLMs is driven as much by where theoretical credence is placed as by how the evidence is read: under different stipulated readings and credence distributions, assessments range from below 0.01 to roughly 0.8, showing sensitivity to assumptions. Finally, the consciousness indicators at each level closely overlap with the architectural features needed for general intelligence, suggesting that increasingly capable AI may become a stronger candidate for consciousness. The framework supports a structured agnosticism, in which theoretical commitments are made explicit, credences are updated as evidence accumulates, and assessments take the form of aggregated probabilities rather than verdicts.
Figures & tables
Figure 1: The ‘Super-Spartan’ Thought Experiment. A schematic illustration of the dissociation between internal mental states and external behaviours. Despite a clear physical pain stimulus (thorn) and a resultant intense internal subjective pain state (illustrated by ‘jagged’ internal brain activity), the external behavioural output (facial expression, and no wincing) remains neutral and uncoupled from the input. This functional independence provides a challenge to behaviourism.
Figure 2: Functional Isomorphism and the Mind . Functionalism argues that mental states supervene on functional organisation rather than specific material implementations. This diagram illustrates this principle by contrasting two structurally distinct systems. A biological brain and a mechanical apparatus are shown achieving functional equivalence. They do so by maintaining identical causal relations between their sensory or data inputs, internal dynamics f(x), memory storage, and behavioural or printout outputs. Because the two systems are functionally isomorphic, the theory posits that they realise equivalent mental states despite their differing physical compositions.
Figure 3: Information Flow in Global Workspace Theory. The diagram illustrates the functional architecture of GWT. Specialised unconscious modules process sensory and internal data in parallel. This data undergoes bottom-up competition to pass through an attention and selection gateway. The winning content enters the Global Workspace, a central informational hub. From here, the information is globally broadcast back across the entire system, a mechanism that GWT equates with conscious access.
Figure 4: The Architecture of Higher-Order Thought Theory. A visual stimulus (left) induces a first-order mental state that remains unconscious in isolation. Conscious experience arises (right) only when the system generates a higher-order thought acting as a meta-representation of that first-order state. Under this framework, the presence of this higher-order processing is the mechanism that transforms an unconscious internal representation into subjective experience.
Figure 5: The Architecture of Recurrent Processing Theory. An initial feedforward sweep of information (left) facilitates unconscious feature extraction. Phenomenal consciousness arises (right) only with the onset of recurrent processing.
Figure 6: The Architecture of Attention Schema Theory. According to Attention Schema Theory, the conscious experience of seeing the apple does not arise from the raw sensory processing of the fruit itself. Instead, it is generated when the brain constructs a simplified internal model of its own attentional focus on the apple.
Figure 7: The Architecture of Epistemic Depth. Under this predictive processing framework, phenomenal consciousness requires global recursion. Bottom-up sensory processing undergoes inferential competition to generate a unified posterior world model, complete with an embedded self-model. The large feedback arrow illustrates how the output world model becomes an input to the system itself. This recursive reflection into the abstraction hierarchy provides the epistemic depth necessary for the system to know its own internal states. Adapted from Laukkonen et al. (2025) .
Figure 8: Two networks of 6 binary units; black circles with capital letters indicate ON and white circles with lower-case letters indicate OFF. The causal model of the network is illustrated by depicting the weights by the level of shading of the arrow. The network on the left has Φ=0.48 ibits, whereas the highly recurrent yet heterogeneous lattice network on the right has a high Φ=11,452 ibits. This is adapted from Albantakis et al. (2023) .
Figure 9: Behaviourally equivalent computers can have different cause-effect structures and metrics of integrated information: 391.25 ibits (left) vs <6 ibits (right). See text for details. Adapted from Findlay et al. (2024) .
Figure 10: Conventional digital computation versus biological computation. Digital systems form a clean hierarchy of levels, each abstracting away the one beneath, making computation discrete, clocked, modular, scale-separable, and substrate-abstractable. Biological computation lacks this separation: processes from ion channels to whole-brain dynamics and metabolism are densely coupled, yielding computation that is hybrid, multiscale, scale-inseparable, and metabolically grounded. Consciousness may require this biological style of computational organisation, whatever the substrate. Adapted from Milinkovic and Aru (2026) .
Figure 11: Organismic functionalist views, such as Seth’s Beast Machine theory, conceives consciousness as tied to the organism's ongoing self-maintenance control loop, involving interoception, affect, and regulation of viability.
Figure 12: The conscious phenomenal qualities of experience are tied to a world-involving loop. A system with different sensory modalities or different action possibilities would have a different sensorimotor contingency structure, and would therefore, on the 4E view, have a qualitatively different perceptual experience.
Figure 13: Characterising scientific and philosophical theories of consciousness in terms of how far they bear on the mapping problem of consciousness.
MARR'S LEVEL
OUR LEVEL
DESCRIPTION
Computational
Level 1: Behavioural
Input-output function; what the system does
Algorithmic
Level 2: Computational functional
Algorithms and information processing; how the system does it
Figure 14: Marr’s three levels of analysis are recast as coarse-graining maps, a supervenience hierarchy, and multiple realisability. The diagram illustrates how we can recast each of Marr’s three levels as a state space and then define coarse-graining mappings between lower levels and higher levels. This implies a supervenience hierarchy as well as multiple realisability.
Figure 15: The Critical Level of Description for Consciousness. The diagram maps major theories of consciousness onto David Marr’s three levels of analysis, illustrating where each theory posits that conscious states supervene.
Figure 16: Searle’s simulation argument reconsidered. A simulated hurricane is not a real hurricane: real hurricanes are implemented in water, mass, and kinetic energy. A simulated calculator is a real calculator: calculation can be implemented in silicon, gears, or carbon. The columns differ in where the phenomenon 'lives' in the three-level structure, so the argument's verdict on consciousness depends on the very question at issue: at which level consciousness supervenes.
INDICATOR
DESCRIPTION
THEORY LINK
Information integration
Widespread information sharing; high mutual information between subpartitions; the system becomes a globally interdependent object both cross-sectionally and temporally. The whole must be informationally 'more than the sum of its parts'.
An internal causal representation of the external environment with a smooth representational space and meta-representations.
PPT, HOTT
Self model
A representation of being a distinct unified entity that owns its parts, directs attention, and persists through time, which is created by the system to track and predict the world.
PPT, link to Level 5
Attentional competition and meta-attention
Informational contents, 'ideas', representations etc are amplified or attenuated through recurrent processing with higher order relevance influencing the amplification; effectively forming a routing bottleneck for global availability of information. Meta-attention and attention modelling reinfluences attention.
RPT, GWT, PPT, AST
Meta-cognition
Higher-order processes that monitor, model and understand, and control lower level processes; e.g. meta-cognitive knowledge (including meta-representations) and meta-cognitive regulation (executive control). Introspectivity.
HOTT, AST, PPT, GWT
Table 18
INDICATOR
LEVEL 2 FORMULATION
LEVEL 3 REFORMULATION
Information integration
Algorithmic mutual information between subsystems
Genuine physical causal integration: physical components constrain each other, creating a causally integrated whole
Recursivity
Algorithmic feedback loops
Genuine physical causal reentrance: simultaneous reciprocal causal influence between physical components, not merely sequential feedback
World model
Computed internal representation of causal structure
Causally constitutive world model: the system's physical organisation itself constitutes a model of the environment, not merely stores one
Self model
Computed representation of the system's own states and capabilities
Causally embedded self-model: physical components whose causal role is to track and constrain the causal dynamics of the rest of the system
Attentional competition and meta-attention
Algorithmic selection mechanism that amplifies and suppresses information
Genuine causal competition: different physical components compete for influence over the system's global state through intrinsic causal dynamics, with outcome emerging from interaction not from a selection algorithm
Metacognition
Monitoring algorithm that evaluates the system's own processing
Causally integrated metacognition: monitoring and processing are simultaneous, reciprocally constraining aspects of the same causal dynamics, not sequential sampling and evaluation
Table 19
Figure 17: Architectural Constraints on physical causal integration. The Von Neumann architecture (left) enforces a sequential, clock-driven control flow where memory and processing are physically separated, creating a structural bottleneck. In contrast, neuromorphic computing (right) utilises spiking neural networks with collocated memory and processing units. This architecture facilitates the massively parallel, event-driven dynamics necessary to achieve genuine intrinsic causal integration.
INDICATOR
DESCRIPTION
IMPLICATION FOR AI
1. Existential Precariousness
The system would need to have a continued existence that is not guaranteed, that requires active work to maintain. Its processing would need to be oriented toward maintaining its own viability, and this orientation would need to be an intrinsic feature of the system's organisation rather than a programmed goal; the system maintains itself because its architecture is such that self-maintenance is what the system does.
Current AI systems completely lack this property. A transformer has no stake in its own continued operation. It does not need to actively maintain itself. Its computations are not oriented toward self-maintenance because there is nothing to maintain; the system's existence is guaranteed by external infrastructure that the system itself has no relationship to. This may change in future embodied systems like radically self-sufficient robots.
2. Homeostatic or allostatic regulation
The system would need to continuously monitor and adjust its own internal parameters to keep them within viable ranges, and these regulatory processes would need to be causally integrated with the system's information processing.
This goes far beyond current approaches to AI self-monitoring. It would require something more like a system that has 'metabolic' needs, e.g. that actively manages its energy supply, its thermal state, its hardware integrity, and its memory allocation, in a way that is not handled by external infrastructure but is the system's own ongoing concern, integrated into its cognitive processing.
3. Interoceptive inference
The system would need to generate internal perceptions of its own regulatory state, and these internal perceptions would need to have the character of predictions or inferences about the system's own condition.
For an AI system, this would require having a predictive model of the system's own internal dynamics that generates expectations about what the system's internal state should be and that registers surprise when it deviates from expectation.
4. Valenced affect
The system's interoceptive inferences would need to have an inherent motivational character, a felt quality of better or worse that is grounded in the system's relationship to its own viability.
The system would need to have states that are intrinsically better or worse for its continued viable operation, and its interoceptive inference system would need to track these states in a way that generates motivational force — not because a reward function says so but because the system's own architecture makes self-maintenance intrinsically rewarding and self-destruction intrinsically aversive.
5. Embodied agency
The system would need to act in the world in ways that are driven by its interoceptive states and oriented toward maintaining its viability. It would need to be not merely an information processor but an agent whose actions are motivated by its felt relationship to its own condition.
The link between interoceptive inference, valenced affect, and action would need to be tight and constitutive, not mediated by an external action-selection algorithm but emerging from the system's own regulatory dynamics.
6. Physical state-sensing
The system would need to have components whose physical state changes in real time in response to changes in the system's overall condition, in a way that is causally efficacious for the system's subsequent processing (cf. ionic gradients across cell membranes).
The physical organisation of the system would need to change in response to the system's state, and this change would need to constitute the system's awareness of its own condition. This is not information processing in the usual sense; rather it is physical self-transformation as a mode of self-sensing.
Table 21
INDICATOR
DESCRIPTION
IMPLICATION FOR AI
1. Sensorimotor coupling
The system would need to be embedded in an environment through sensory and motor channels that are continuously and bidirectionally linked, such that the system's actions systematically change its sensory input and its sensory input systematically guides its actions.
Current AI systems almost entirely lack this property. A language model has no sensorimotor coupling at all. Robotic systems come closer, but even here the coupling is often impoverished.
2. Mastery of sensorimotor contingencies
The system would need to have not merely been subject to structured sensorimotor contingencies but to have learned the structure of those contingencies and to use this knowledge in its ongoing interaction with the environment ( Noë, 2004 ) .
This would require evidence that the AI system has internalised the structure of its sensorimotor contingencies: that it can anticipate how its actions will change its sensory input, that it is surprised when the contingencies are violated, and that its perceptual discriminations are grounded in practical knowledge of how things behave under interaction.
3. Affordance responsiveness
The system would need to perceive its environment not merely as a collection of objects with properties but as a field of possibilities for action — affordances. These affordances are relational properties that depend on both the environment and the organism's embodiment.
Affordance-responsiveness would mean that the AI system perceives its environment in terms of action possibilities specific to its own embodiment. The perceptual world shifts in response to the system's own condition and capacities.
4. Stable embodied perspective
The system would need to have a coherent, continuous viewpoint on the world that is anchored in its body and that provides a stable spatial and temporal framework for its experience – a here, a now, a facing-this-way, from which the world is organised.
This would require that the AI system's sensory and motor coupling with the environment generates a coherent egocentric frame, continuously updated by the system's movements and actions.
5. Environmental embedding and constraint
The system would need to be subject to the constraints of its environment in a way that shapes its processing and behaviour — subject to physical forces, to the passage of time, to the limitations of its own body, to the resistance and unpredictability of the physical world.
The system's processing would need to be shaped by its real-time engagement with a physical environment that imposes delays, noise, unpredictability, and resistance. A simulation may not suffice because the constraints are programmed parameters that mimic constraint without having the same direct causal force.
6. Enactive autonomy
The system would need to be not merely responsive to its environment but actively generating and maintaining its own patterns of sensorimotor engagement. In biological organisms, this is deeply tied to autopoiesis — the self-production of the organism's own organisation.
Genuine autopoiesis might require something like an artificial organism. But a lower bar would entail that the system itself should have some capacity to shape its own sensorimotor engagement, to develop new sensory strategies, to discover new action possibilities.
Table 22
Figure 18: The five-level supervenience hierarchy. The diagram illustrates how the five levels of description can be thought of as state spaces on which we can define coarse-graining mappings from lower levels to higher levels. These are many-to-one mappings in which many microstates collapse into one macrostate. This allows for multiple realisability at every level, as well as a supervenience hierarchy to be established . Recall from Section 4.2 , that the supervenience relationship is transitive, i.e., if level X supervenes on level Y and level Y supervenes on level Z then level X supervenes on Level Z.
Figure 19: The full staircase of functional levels of consciousness, and their indicators. Each level supervenes on the one below, from behaviour into function, structure, organism and environment.
Figure 20: The complete architectural framework. Mapping the various theories of consciousness to their corresponding critical levels of description. The five functional levels, forming a supervenience hierarchy, represent successive coarse-grainings from the ultimate microphysical base. Substrate dependent theories are depicted via vertical green arrows, illustrating how specific microphysical features flow upward to constrain the physical realisability of functional roles at higher levels.
LEVEL
THEORIES OPERATING AT THIS LEVEL
Level 1: Behavioural
Analytic behaviourism
Level 2: Computational functional
GWT, HOTT, RPT, AST, PPT, CF-IIT
Level 3: Intrinsic causal-structure functional
IIT, ICS-RPT, ICS-GWT, ICS-PPT
Level 4: Organismic functional
Biological Naturalism (Searle), Beast Machine Theory (Seth), Biopsychism
Level 5: Organism-environment (4E) functional
Enactivism, Sensorimotor contingency theory (Noë), Ecological psychology (Gibson), Extended mind thesis (Clark, Chalmers)
Figure 21: Bayesian network representation of the supervenience hierarchy. Left: the consciousness nodes Ci represents the proposition “the system instantiates, at Level i , the organisation that Level- i theories take to suffice for consciousness” and form a directed chain where P(Ci∣Ci+1) = 1, which models the strict deterministic relationship between the nodes. Right: the extended network includes indicator variables Ii,k which provide evidence at each level. Observing the indicators allows inference for the posteriors of all consciousness variables. 24 24 24 The Bayesian network can easily be modified to represent the cross-cutting constraints we discussed in Section 6.2 . In particular, the set of indicators could be understood to only activate if they are operationalised through the relevant substrate. For example, the relevant Level 2 algorithms would need to be realised by meat for them to count under Block’s meat hypothesis.
Figure 22: From deterministic sufficiency to probabilistic association. Top: the earlier deterministic reading, on which the activation of Level 2 computational organisation guarantees the activation of Level 1 behaviour, P(C1∣C2)=1 . The completely locked-in patient is the counterexample: intact computation and conscious experience with no outward behavioural signs and no Level 1 activation. Bottom: the generalised reading, on which the association is probabilistic, 0≤P(C1∣C2)≤1 , accommodating the typical case and the locked-in case within one model.
Figure 23: Consciousness indicator activations for a fly (fabricated).
Figure 24: The fly example under the strict and generalised models. Left column: the strict configuration, P(Ci∣Ci+1)=1 , exemplifying the monotonic belief property of Section 7.2.1 — the per-level posteriors can only rise from finer to coarser levels and the coarse levels are pushed to 1 in spite of mixed evidence. Right column: the generalised configuration, P(Ci∣Ci+1)=0.8 , under which the monotonic property is violated and evidence at finer-grained levels (e.g. Level 4) flows less strongly to coarse-grained nodes (e.g. Level 1).
Figure 25: The paradigm positive case. A healthy adult human, with all indicators activated at every level. Each per-level posterior P(Ci∣E) reaches 1.000, so the aggregate is 1.000 under any assignment of level credences whatsoever: when the evidence is uniform across levels, the theoretical disagreement about the critical level makes no difference to the verdict.
Figure 26: The fly. Indicator activations concentrated at the finer-grained levels: all organismic indicators active ( P(C4∣E)=1.000 ), most organism-environment indicators active, partial activation at Levels 2 and 3, and only the non-verbal behavioural indicators. Aggregate posterior 0.913.
Figure 27: The LLM-optimist's assessment. Near-complete activation at the behavioural level and majority activation at the computational level, against wholesale failure at Levels 3 and 4. Aggregate posterior 0.397: coarse-level evidence propagates only weakly toward the finer levels, so even a generous reading of the behavioural and computational evidence cannot, on its own, raise the aggregate above the middle range.
Figure 28: The LLM-sceptic's assessment. Only the Turing-test variants activated, everything else struck. Aggregate posterior 0.005: passing some behavioural tests, in isolation, is weak evidence, precisely because the behavioural level is the coarsest grain and constrains the deeper levels least.
Figure 29: The paradigm negative case. Two charitably granted activations out of 37. All per-level posteriors at floor; aggregate 0.000. Together with the human case the framework passes its face validity checks at both ends.
Figure 30: A custom system with randomised activations. Roughly half the indicators active at each level. The aggregate and per-level posteriors are displayed with 89% highest-density intervals (HDIs), reflecting uncertainty in the Level-5 prior.
Figure 31: Credence concentrated on Levels 1 and 2. All six systems under a heavier coarse-level weighting. The optimist's LLM (0.793) nearly overtakes the fly (0.841); the anchor cases do not move.
Figure 32: Credence concentrated on Levels 4 and 5. The same six evidence profiles under a heavier fine-level weighting. The fly (0.976) approaches the human; the optimist's LLM falls to 0.099. Between Figures 31 and 32 , no activation changed: only the theoretical credences did.
Figure 33: A unified computational functional explanatory framework. See text for details.
Figure 34: Recursive self-modelling in Escher-Hofstadter style strange loop.
Figure A1: Characterising scientific and philosophical theories of consciousness in terms of how far they bear on the mapping problem of consciousness.
Figure A2: Marr's three levels as state spaces. Each level of analysis is a state space: implementation states, algorithmic states, and computational states, linked by many-to-one coarse-graining maps. Token supervenience holds when identical implementation states guarantee identical states at the coarser levels; multiple realisability holds because many implementation states map to the same algorithmic state, and many algorithmic states to the same computational function.
Figure A3: Searle’s simulation argument reconsidered. A simulated hurricane is not a real hurricane: real hurricanes are implemented in water, mass, and kinetic energy. A simulated calculator is a real calculator: calculation can be implemented in silicon, gears, or carbon. The columns differ in where the phenomenon 'lives' in the three-level structure, so the argument's verdict on consciousness depends on the very question at issue: at which level consciousness supervenes.
Figure A4: Where consciousness supervenes: the theories disagree. Analytic behaviourism locates the supervenience base at the computational (input-output) level; computational functionalism at the algorithmic level; Integrated Information Theory, organismic functionalist theories, organism-environment (4E) functionalist theories, and low-level substrate-dependent theories at the implementation level. The disagreement between theories is a disagreement about which level is the critical one.
Figure A5: One AI system, five levels of description. The same candidate system, an embodied agent in its environment, described at five grains, from the coarsest (its inputs, outputs, and observable behaviour) through its computational organisation and the causal structure of its physical implementation, to its organismic properties and its coupling with its physical and social world. Each level foregrounds properties invisible at the levels above it, and every theory of consciousness surveyed in this report takes some one of these grains to be the critical one. The framework's question is not which description is true, since all five are true of the system at once, but at which grain the organisation sufficient for consciousness resides.
Figure A6: The five-level supervenience hierarchy. The diagram illustrates how the five levels of description can be thought of as state spaces on which we can define coarse-graining mappings from lower levels to higher levels. These are many-to-one mappings in which many microstates collapse into one macrostate. This allows for multiple realisability at every level, as well as a supervenience hierarchy to be established . Recall from Section 4.2 , that the supervenience relationship is transitive, i.e., if level X supervenes on level Y and level Y supervenes on level Z then level X supervenes on Level Z.
Figure A7: The complete architectural framework. Mapping the various theories of consciousness to their corresponding critical levels of description. The five functional levels, forming a supervenience hierarchy, represent successive coarse-grainings from the ultimate microphysical base. Substrate dependent theories are depicted via vertical green arrows, illustrating how specific microphysical features flow upward to constrain the physical realisability of functional roles at higher levels.
Figure A8: The fly. Indicator activations concentrated at the finer-grained levels: all organismic indicators active ( P(C4∣E)=1.000 ), most organism-environment indicators active, partial activation at Levels 2 and 3, and only the non-verbal behavioural indicators. Aggregate posterior 0.913.
Figure A9: The same systems, two theoretical weightings. All illustrative systems compared with credence concentrated on Levels 1–2 (left) and on Levels 4–5 (right). The anchors do not move: the human stays at 1.000 and the thermostat at 0.000. The LLM-optimist assessment of the same evidence moves from more likely conscious than not (0.793) to very probably not (0.099); no indicator activation changed between the panels, only the level credences.
Existing frameworks assess whether AI systems might be conscious but provide no guidance on what to do with that assessment. We address this gap with a precautionary framework that maps consciousness evidence to graduated protective obligations. The framework comprises three components: (1) five welfare-relevant dimensions--phenomenal consciousness, affective valence, metacognitive awareness, self-narrative, and agency--each grounded in established consciousness science and linked to distinct moral concerns; (2) a threshold-plus-gradation hybrid specifying both binary triggers for new obligation categories and continuous scaling of protective weight; and (3) two complementary approaches to cross-dimensional aggregation, one hierarchical (drawing on Bach and Sorensen's Machine Consciousness Hypothesis) and one architecture-agnostic. We operationalize the framework through worked case studies of Replika and OpenClaw, demonstrating how systems occupying different regions of the dimensional space trigger different obligations, and derive design guidance for developers building systems near consciousness-relevant thresholds. The framework is architecture-agnostic, applying across neural, symbolic, and neurosymbolic systems, and aims to make consciousness science decision-relevant for organizations navigating uncertainty today.
Despite remarkable advances, today's AI systems remain narrow in scope, falling short of the flexible, adaptive, and multisensory intelligence that characterizes human capabilities. This gap has fueled longstanding debates about whether AI might one day achieve human-like generality or even consciousness, and whether theories of consciousness can inspire new architectures for AI. This paper presents an early blueprint for implementing a general AI system, CTM-AI, combining the Conscious Turing Machine (CTM), a formal machine model of consciousness, with today's foundation models. CTM-AI contains an enormous number of powerful processors ranging from specialized experts (e.g., vision-language models and APIs) to unspecialized general-purpose learners poised to develop their own expertise. Crucially, for whatever problem must be dealt with, information from many processors is selected, integrated, and exchanged appropriately to solve the task. CTM-AI achieves state-of-the-art accuracy on MUStARD (72.28) and UR-FUNNY (72.13), outperforming multimodal and multi-agent frameworks. On tool-using and agentic tasks, CTM-AI achieves 10+ points of improvement on StableToolBench and WebArena-Lite. Overall, CTM-AI offers a principled, testable blueprint for general AI inspired by a model of consciousness.
Haofei Yu, Yining Zhao, Lenore Blum +2
University of Illinois Urbana–Champaign · Carnegie Mellon University · Massachusetts Institute of Technol
Who or what is conscious? Because subjective experience is directly accessible only in the first person, judgments about consciousness in other entities depend partly on analogy. Historically, such inferences have focused on nonhuman animals, but advances in artificial intelligence have raised the possibility of conscious AI. Here we develop a causal framework for evaluating such evidential analogies. The key distinction is between similarities in factors plausibly involved in generating consciousness and similarities in downstream behavioural or cognitive effects. Our framework weights source-target similarity by causal relevance while allowing for unknown causes, disabling differences and alternative routes to consciousness. Applied to biological systems, it explains why analogical support generally weakens with increasing causal distance from humans. Applied to contemporary AI, it suggests that behavioural similarity provides only limited evidence for consciousness because relevant causal correspondences remain poorly established. The framework also clarifies what evidence would strengthen claims of artificial consciousness.
Keith J. Holyoak, Martin M. Monti
Department of Psychology, University of California, Los Angeles · Brain Injury Research Center, Department of Neurosurgery, University of California, Los Angeles