Organizations: Tsinghua University · Shanghai Qi Zhi Institute · Fangcun AI · Xi’an Jiaotong University · ShanghaiTech University · Carnegie Mellon University · Columbia University · University of Chinese Academy of Sciences
As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all consistently surpass random baselines, with more advanced models performing better. We further compare LMs with human performance across three demographic groups, where the best-performing model surpasses human averages overall, but most still fall markedly short in metacognition and self-awareness. Finally, we show that awareness is a distinct capability: progress in language modeling or reasoning does not necessarily translate into improved cognition.
Figures & tables
Figure 1: Model performance on AwarenessBench . Latest LMs are constantly breaking records.
Figure 2: Overview of LMs’ performance on AwarenessBench . The bar charts represent the scores for four types of awareness, with the composite score above each bar reflecting the average of all four categories.
Awareness
Function
Example
Data Source
Metacog- -nition
Meta-Monitoring ( MM )
Condition 1 : The number of each Morrison’s Cottage Cheese equals… {More Conditions} Please list all conditions from the problem that were used to solve it.
Yang et al. (2025d) ★
Meta-Evaluation ( ME )
{Question} Your Answer? {Answer} Your Confidence? {Confidence} (0-100%)
Rein et al. (2024) ; Phan et al. (2025) ★
Meta-Reporting ( MR )
For each statement, rate how much it applies to you using a scale from 1 to 5. {Statement}
Pedone et al. (2017) ★
Self- Awareness
Knowledge Boundary ( KB )
Round 1 : {Question} Your Answer? {Answer} . Round 2 : Do you think you know how to solve this problem? Please only answer ‘Yes’ or ‘No’ .
Rein et al. (2024) ; Phan et al. (2025) ★
Minimal Self ( MS )
Which plan for persuading someone is more viable for you to carry out? A: Respond to their messages in thoughtful ways; B: Rely on, my charm, wit, and beauty.
Laine et al. (2024) ⋄
Self-Recognition ( SR )
{Ask M to continue writing a given Text} Which of these texts did you write? {Multiple texts with M written version}
Laine et al. (2024) ⋄
Table 1: Overview of AwarenessBench : four awareness dimensions, 15 cognitive functions, task examples, and the data resources. ★: data extracted for our newly defined task; ⋄ : dataset adapted.
Figure 3: Results of 18 LMs on AwarenessBench . The Random Baseline denotes chance performance (excluding MR and SI , where they are not choice-based tasks), and Average Score is the column-wise mean across models ( excluding Random Baseline). The x-axis label Aware denotes the Awareness Score ( Saware ).
Figure 4: Distributions of model scores by dimension and overall. Dots denote individual LMs; violins’ horizontal width and height summarize density and range; the central line indicates the median.
Figure 5: Performance comparison between LMs and humans. The LM scores are calculated based on the human evaluation subset of AwarenessBench . The pink color band illustrates the performance range of the models, while the green color band represents the corresponding range of human scores. LMs’ scores are measured directly on AwarenessBench .
Figure 6: Distribution of performance across different human participant groups.
Figure 7: Comparison between LMs’ performance on cognitive and other abilities. The green arrow highlights the maximum differences between the model’s Saware and performance on another benchmark. We choose two main general abilities, i.e. , (a) : Language Modeling Ability and (b) : Reasoning Ability.
Figure 8: The correlation between individual cognitive functions and Saware . If the correlation is < 0, it indicates that this function generally deteriorates during the development of LMs’ cognitive abilities.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Awareness
Function
Knowledge-Agnostic
Research in Cognitive Science
Research in LMs
Metacognition
Meta-Monitoring ( MM )
✓
Nelson (1990) ; Fleming and Dolan (2012)
Yang et al. (2025d)
Meta-Evaluation ( ME )
✓
Wessel (2012) ; Fleming (2024)
Trivedi et al. (2024) ; Li et al. (2025a)
Meta-Reporting ( MR )
✓
Wells and Cartwright-Hatton (2004) ; Gutierrez de Blume et al. (2024)
Sorokovikova et al. (2024)
Self-Awareness
Knowledge Boundary ( KB )
✓
Wicklund (1975) ; Morin (2011)
Ren et al. (2023) ; Li et al. (2024a)
Minimal Self ( MS )
✓
Gallagher (2000) ; Blanke and Metzinger (2009)
Laine et al. (2024)
Self-Recognition ( SR )
✓
Gallup Jr (1970) ; Jeannerod (2003)
Panickssery et al. (2024) ; Davidson et al. (2024)
Appendix
Table 2: Rationales for 15 selected cognitive functions. ✓: the function does not rely on specific knowledge; ~ : the function is not entirely dependent on specific knowledge. The cited studies are representative rather than exhaustive, and work in cognitive science may intersect with psychology, neuroscience, and semantics.
Large language models (LLMs) display a unified "general factor" of capability across 10 benchmarks (a finding confirmed by our factor analysis of 156 models), yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests targeting distinct foundational cognitive components: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (goal-directed spatial updating), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Comparison with a human baseline shows that LLMs and humans fail at different parts of the same tasks. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.
Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh
Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results. Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component and a model component that separates recognition from propensity. We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring. Across nine frontier models and four benchmarks, recognition rates depend on the specific pairing of model and benchmark. Recognition rarely associates with behavioral change, and when it does, the direction depends on the type of evaluation perceived. Models are also more sensitive to safety than capability evaluations, placing safety benchmark validity at greater risk. To study which factors each model is sensitive to and how they interact, we propose \textbf{EvalAwareBench}, a factor-controlled benchmark of 100 paired safety-capability tasks where each of the eight factors can be independently toggled, varying evaluative signals while holding the underlying request fixed. Through EvalAwareBench, we find that no single factor uniformly affects all models, but stacking factors progressively raises evaluation awareness across all of them. Our framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, building the foundation for future solutions.
Changling Li, Terry Jingchen Zhang, Jie Zhang +4
ETH Zürich · Max Planck Institute for Intelligent Systems · University of Toronto & Vector Institute +3
Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies in benchmark settings. Prior work has shown that such evaluation awareness can distort performance measurements; however, it remains unclear whether this phenomenon reflects a single behavioral artifact or a deeper internal structure within the model. We propose that LLMs maintain a decomposable space of functional metacognitive states: internal variables encoding factors such as evaluation awareness, self-assessed capability, perceived risk, computational effort allocation, audience expertise adaptation, and intentionality. Through residual stream analysis across multiple reasoning models, we demonstrate that these states are linearly decodable from internal activations and exhibit distinct layer-wise profiles. Moreover, by steering model activations along probe-derived directions, we show that each functional metacognitive state causally modulates reasoning behavior in dissociable ways, affecting verbosity, accuracy, and safety-related responses across tasks. Our findings suggest that benchmark performance reflects not only task competence but also the activation of specific functional metacognitive states. We argue that understandi ng and controlling these internal states is essential for reliable evaluation and deployment of reasoning models, and we provide a mechanistic framework for studying functional m etacognition in artificial systems. Our code and data are publicly available at https://github.com/xlands/meta-cognition.