Organizations: Tsinghua University · Shanghai Qi Zhi Institute · Fangcun AI · Xi’an Jiaotong University · ShanghaiTech University · Carnegie Mellon University · Columbia University · University of Chinese Academy of Sciences
As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all consistently surpass random baselines, with more advanced models performing better. We further compare LMs with human performance across three demographic groups, where the best-performing model surpasses human averages overall, but most still fall markedly short in metacognition and self-awareness. Finally, we show that awareness is a distinct capability: progress in language modeling or reasoning does not necessarily translate into improved cognition.
Figures & tables
Figure 1: Model performance on AwarenessBench . Latest LMs are constantly breaking records.
Figure 2: Overview of LMs’ performance on AwarenessBench . The bar charts represent the scores for four types of awareness, with the composite score above each bar reflecting the average of all four categories.
Awareness
Function
Example
Data Source
Metacog- -nition
Meta-Monitoring ( MM )
Condition 1 : The number of each Morrison’s Cottage Cheese equals… {More Conditions} Please list all conditions from the problem that were used to solve it.
Yang et al. (2025d) ★
Meta-Evaluation ( ME )
{Question} Your Answer? {Answer} Your Confidence? {Confidence} (0-100%)
Rein et al. (2024) ; Phan et al. (2025) ★
Meta-Reporting ( MR )
For each statement, rate how much it applies to you using a scale from 1 to 5. {Statement}
Pedone et al. (2017) ★
Self- Awareness
Knowledge Boundary ( KB )
Round 1 : {Question} Your Answer? {Answer} . Round 2 : Do you think you know how to solve this problem? Please only answer ‘Yes’ or ‘No’ .
Rein et al. (2024) ; Phan et al. (2025) ★
Minimal Self ( MS )
Which plan for persuading someone is more viable for you to carry out? A: Respond to their messages in thoughtful ways; B: Rely on, my charm, wit, and beauty.
Laine et al. (2024) ⋄
Self-Recognition ( SR )
{Ask M to continue writing a given Text} Which of these texts did you write? {Multiple texts with M written version}
Laine et al. (2024) ⋄
Table 1: Overview of AwarenessBench : four awareness dimensions, 15 cognitive functions, task examples, and the data resources. ★: data extracted for our newly defined task; ⋄ : dataset adapted.
Figure 3: Results of 18 LMs on AwarenessBench . The Random Baseline denotes chance performance (excluding MR and SI , where they are not choice-based tasks), and Average Score is the column-wise mean across models ( excluding Random Baseline). The x-axis label Aware denotes the Awareness Score ( Saware ).
Figure 4: Distributions of model scores by dimension and overall. Dots denote individual LMs; violins’ horizontal width and height summarize density and range; the central line indicates the median.
Figure 5: Performance comparison between LMs and humans. The LM scores are calculated based on the human evaluation subset of AwarenessBench . The pink color band illustrates the performance range of the models, while the green color band represents the corresponding range of human scores. LMs’ scores are measured directly on AwarenessBench .
Figure 6: Distribution of performance across different human participant groups.
Figure 7: Comparison between LMs’ performance on cognitive and other abilities. The green arrow highlights the maximum differences between the model’s Saware and performance on another benchmark. We choose two main general abilities, i.e. , (a) : Language Modeling Ability and (b) : Reasoning Ability.
Figure 8: The correlation between individual cognitive functions and Saware . If the correlation is < 0, it indicates that this function generally deteriorates during the development of LMs’ cognitive abilities.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Awareness
Function
Knowledge-Agnostic
Research in Cognitive Science
Research in LMs
Metacognition
Meta-Monitoring ( MM )
✓
Nelson (1990) ; Fleming and Dolan (2012)
Yang et al. (2025d)
Meta-Evaluation ( ME )
✓
Wessel (2012) ; Fleming (2024)
Trivedi et al. (2024) ; Li et al. (2025a)
Meta-Reporting ( MR )
✓
Wells and Cartwright-Hatton (2004) ; Gutierrez de Blume et al. (2024)
Sorokovikova et al. (2024)
Self-Awareness
Knowledge Boundary ( KB )
✓
Wicklund (1975) ; Morin (2011)
Ren et al. (2023) ; Li et al. (2024a)
Minimal Self ( MS )
✓
Gallagher (2000) ; Blanke and Metzinger (2009)
Laine et al. (2024)
Self-Recognition ( SR )
✓
Gallup Jr (1970) ; Jeannerod (2003)
Panickssery et al. (2024) ; Davidson et al. (2024)
Appendix
Table 2: Rationales for 15 selected cognitive functions. ✓: the function does not rely on specific knowledge; ~ : the function is not entirely dependent on specific knowledge. The cited studies are representative rather than exhaustive, and work in cognitive science may intersect with psychology, neuroscience, and semantics.