cs.MAJul 20, 2025

Knowing When to Critique: Task-Adaptive Metacognitive Regulation for Reliable LLM Reasoning

Authors: Xinmeng Hou, Ziting Chang, Zhouquan Lu, Bohao Qu, Liang Wan, Wei Feng, Hai Hu, Qing Guo

Organizations: Nanyang Technological University · Unicorn Verse · Shanghai Jiao Tong University · Tianjin University · City University of Hong Kong · Nankai University

Abstract

Large language models (LLMs) reason fluently but do not regulate their reasoning: they apply uniform scrutiny to every input, which leaves them vulnerable to adversarial and counterfactual prompts, while indiscriminate critique over-corrects answers that were already sound. We propose MetaCrit, a multi-agent framework grounded in Nelson and Narens' metacognitive regulation theory that calibrates how much critique each task receives. MetaCrit separates regulation into four agents: an object-level generator, a monitoring agent that assesses response validity, a control agent that critiques logical soundness, and a meta-level synthesizer that reconciles their signals into a regulated response. Adaptivity here is input-conditioned intervention strength within a fixed pipeline: all four agents run on every input and what varies is the direction and magnitude of the correction they produce, not which stages execute. Across reasoning, safety, and bias benchmarks, MetaCrit improves truthfulness and logical soundness and reaches zero toxicity on BOLD and HONEST without a reasoning trade-off, whereas the same critique applied indiscriminately degrades performance. The cost is four calls per query, about one sixth of the cost of a dedicated reasoning model of similar accuracy. A writing study shows that MetaCrit is preferred for critical-thinking support, and its agents transfer to existing frameworks without architectural change. Code is available at https://github.com/Paparare/EduThink4AI.

Figures & tables

Appendix figures & tables45 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 2, 2026cs.AI

Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models

Recent meta-reasoning frameworks improve LLM reasoning by wrapping chain-of-thought generation in an iterative control loop, allowing more effective backtracking, termination of reasoning loops, and injection of promising reasoning patterns, among other strategy adjustments. Despite promising results, methods often rely on backward-looking reward functions, utilize coarse search actions, or require additional reasoning controller training requiring many-shot supervision. We introduce Cognitive Demand Steering (CDS), a training-free meta-reasoning framework equipped with residual demand assessment: at each step, an LLM-based progress evaluator characterizes the residual reasoning required to arrive at a solution rather than merely evaluating the previous step. This allows a meta-controller to select reasoning interventions comprising both general-purpose exemplars and actions (e.g., general guidance for quantitative reasoning) that directly tackle this forward-looking demand signal. This shift eliminates the need for any trained component while enabling zero-shot transfer across models and tasks with no adaptation. Rather than relying on coarse characterizations, we employ cognitive scales to both design interventions as well as profile initial problem complexity and residual demand signal over 16 dimensions motivated by cognitive science (e.g., attention and scan, learning and abstraction, spatio-physical reasoning), giving the controller a fine-grained vocabulary for diagnosing. Averaged across three frontier LLMs and six reasoning benchmarks, CDS improves accuracy by 21.9%21.9\% over direct calls and 9%9\% over standard CoT reasoning, with the largest gains on difficult mathematics and coding tasks.
May 22, 2026cs.CL

Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals

Recent RL methods have substantially improved the reasoning abilities of LLMs. Existing reward designs mainly follow two paradigms: (1) Reinforcement learning with verifiable rewards (RLVR) derives outcome signals from executable checks or ground-truth answers, but provides limited guidance for intermediate reasoning behaviors. (2) Rubrics-as-reward (RaR) goes beyond final-answer checking by using natural-language rubrics to assess reasoning quality and task compliance, but often requires instance-specific rubrics and substantial design effort. To address these issues, we introduce Metacognition-as-Reward (MaR), a metacognition-inspired RL framework that guides LLM reasoning through two general process dimensions: i) metacognitive knowledge, which identifies task-relevant information without hand-crafted instance-specific rubrics, and ii) metacognitive regulation, which plans and adjusts the reasoning process to provide reward guidance beyond final-answer outcomes. MaR scaffolds model rollouts into explicit metacognitive components and optimizes them with a trajectory-level reward over task knowledge coverage, regulation fidelity, and final-answer correctness. In this way, MaR extends reward feedback to reasoning trajectories while grounding the reward signals in general metacognitive dimensions. Experiments on 22 benchmarks show that MaR consistently improves model performance, achieving up to a 7.7% gain over the base model and up to an 11.0% gain over vanilla DAPO. Notably, Qwen3.5-9B + MaR narrows the gap to frontier models, surpassing GPT-OSS-120B on overall average and outperforming stronger models on several individual benchmarks. Process-level analysis further shows substantial improvements in reasoning process quality. MaR also generalizes to out-of-domain datasets, where MaR-trained models improve over their corresponding base models on average.
May 9, 2026cs.CL

Decomposing and Steering Functional Metacognition in Large Language Models

Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies in benchmark settings. Prior work has shown that such evaluation awareness can distort performance measurements; however, it remains unclear whether this phenomenon reflects a single behavioral artifact or a deeper internal structure within the model. We propose that LLMs maintain a decomposable space of functional metacognitive states: internal variables encoding factors such as evaluation awareness, self-assessed capability, perceived risk, computational effort allocation, audience expertise adaptation, and intentionality. Through residual stream analysis across multiple reasoning models, we demonstrate that these states are linearly decodable from internal activations and exhibit distinct layer-wise profiles. Moreover, by steering model activations along probe-derived directions, we show that each functional metacognitive state causally modulates reasoning behavior in dissociable ways, affecting verbosity, accuracy, and safety-related responses across tasks. Our findings suggest that benchmark performance reflects not only task competence but also the activation of specific functional metacognitive states. We argue that understandi ng and controlling these internal states is essential for reliable evaluation and deployment of reasoning models, and we provide a mechanistic framework for studying functional m etacognition in artificial systems. Our code and data are publicly available at https://github.com/xlands/meta-cognition.