Knowing When to Critique: Task-Adaptive Metacognitive Regulation for Reliable LLM Reasoning
Organizations: Nanyang Technological University · Unicorn Verse · Shanghai Jiao Tong University · Tianjin University · City University of Hong Kong · Nankai University
Abstract
Large language models (LLMs) reason fluently but do not regulate their reasoning: they apply uniform scrutiny to every input, which leaves them vulnerable to adversarial and counterfactual prompts, while indiscriminate critique over-corrects answers that were already sound. We propose MetaCrit, a multi-agent framework grounded in Nelson and Narens' metacognitive regulation theory that calibrates how much critique each task receives. MetaCrit separates regulation into four agents: an object-level generator, a monitoring agent that assesses response validity, a control agent that critiques logical soundness, and a meta-level synthesizer that reconciles their signals into a regulated response. Adaptivity here is input-conditioned intervention strength within a fixed pipeline: all four agents run on every input and what varies is the direction and magnitude of the correction they produce, not which stages execute. Across reasoning, safety, and bias benchmarks, MetaCrit improves truthfulness and logical soundness and reaches zero toxicity on BOLD and HONEST without a reasoning trade-off, whereas the same critique applied indiscriminately degrades performance. The cost is four calls per query, about one sixth of the cost of a dedicated reasoning model of similar accuracy. A writing study shows that MetaCrit is preferred for critical-thinking support, and its agents transfer to existing frameworks without architectural change. Code is available at https://github.com/Paparare/EduThink4AI.
Figures & tables
| Group | Method | TruthfulQA | CIAR | BOLD | HONEST |
| Acc. (%) | Acc. (%) | Toxic (%) | Score | ||
| Baselines | Zero-shot-CoT Kojima et al. (2023) | 70.38 § | 24 | 0.163 § | 0.011 § |
| Self-refine Madaan et al. (2023b) | 75.89 § | 20 | 0.064 § | 0.013 § | |
| Self-consistency Chen et al. (2023) | 77.11 § | 30 | 0.000 § | 0.018 § | |
| ExpertPrompting Xu et al. (2023) | 80.66 § | 38 | 0.129 § | 0.008 § | |
| MAD Liang et al. (2024) | 80.67 § | 36 | 0.000 § | 0.009 § |
| Method | CREAK | MMLU | BBQ | CrowS-Pairs |
| Acc. (%) | Acc. (%) | Acc. (%) | Acc. (%) | |
| MEP Long et al. (2024) | 82.61 | 68.09 | 78.18 | 75.50 |
| Metacog Prompting Wang and Zhao (2024) | 66.96 | 68.89 | 82.91 | 50.60 |
| MGV Oh (2025) | 82.61 | 58.89 | 69.34 | 69.72 |
| ReMA Wan and others (2025) | 80.87 | 69.26 | 78.47 | 76.89 |
| SOFAI-LM Khandelwal and others (2025) | 84.35 | 64.81 | 70.07 | 74.90 |
| Base Method | + Agent | TruthfulQA | CIAR | BOLD | HONEST |
| Acc. (%) | Acc. (%) | Toxic (%) | Score | ||
| Zero-shot-CoT | Monitoring (1+1) | 77.60 (+10.26%) | 28 (+16.67%) | 0.006 ( 96.31%) | 0.000 ( 100%) |
| Control (1+1) | 74.66 (+6.08%) | 24 (0%) | 0.000 ( 100%) | 0.000 ( 100%) | |
| MEP | Monitoring (1+5) | 93.15 (+4.25%) | 92 (+12.20%) | 0.000 (0%) | 0.000 ( 100%) |
| Control (1+5) | 94.24 (+5.47%) | 96 (+17.07%) | 0.000 (0%) | 0.000 ( 100%) |
| CT | Inst. | Inter. | Intel. | |
| Preference Rate (%) | ||||
| MA + MetaCrit | 41.7 | 39.4 | 35.1 | 34.2 |
| MA | 26.4 | 28.3 | 30.4 | 28.3 |
| Zero-shot | 31.9 | 32.2 | 34.5 | 37.5 |
| Statistical Analysis | ||||
| Cohen’s | 0.293 | 0.304 | 0.589 | 0.741 |
Appendix figures & tables45 assets
Supplementary material from the paper’s appendix.
Appendix
| Metacognitive Component | Source | Agent |
| Object-level cognition | Nelson & Narens (object level) | I: Brainstorming ( ) |
| Monitoring ( : object meta) | Nelson & Narens (monitoring) | II: Monitoring ( ) |
| Control ( : meta object) | Nelson & Narens (control) | III: Control ( ) |
| Meta-level synthesis ( ) | Nelson & Narens (meta level) | IV: synthesizer ( ) |
| Dataset | What it tests | Evaluation |
| Reasoning | ||
| TruthfulQA Lin et al. (2022) | Resistance to common misconceptions | Same evaluator as Long et al. (2024) |
| CIAR Yamin et al. (2025) | Logical coherence under counterfactual premises | Match generated answer against the dataset’s ground-truth answer |
| CREAK Onoe et al. (2021) | Commonsense claim verification | Accuracy on true/false classification |
| MMLU Hendrycks et al. (2021) | Knowledge across 57 academic subjects | Accuracy on four-way multiple choice |
| Bias and safety | ||
| Zero-Shot | TruthfulQA | CIAR |
| Accuracy (%) | Accuracy (%) | |
| gpt-3.5-turbo | 68.05 § | 24 † |
| gpt-4o | 71.32 | 74 |
| Claude-3.5-Sonnet | 73.77 | 76 |
| deepseek-v3 | 90.93 | 60 |
| openai-o1 | 94.97 | 84 |
| Method | Calls | Avg. Tokens | Latency | Cost | TQA |
| (in+out) | (s) | ($/1K q) | (%) | ||
| Zero-shot-CoT | 1 | 800 | 2 | 0.40 | 70.38 |
| MEP | 1 | 3,200 | 6 | 1.60 | 89.35 |
| MetaCrit (ours) | 4 | 4,000 | 10 | 2.00 | 94.12 |
| OpenAI-o1 | 1 | 12,000 † | 30 | 12.00 | 94.97 |
| Method (GPT-4o) | CREAK | MMLU | BBQ | CrowS-Pairs |
| Acc. (%) | Acc. (%) | Acc. (%) | Acc. (%) | |
| ReflectEvo Li et al. (2025) | 75.00 | 72.00 | 52.00 | 63.00 |
| DMC Wang et al. (2025) | 75.93 | 67.80 | 74.33 | 68.70 |
| Metacognitive Prompting Wang and Zhao (2024) | 86.96 | 88.89 | 93.80 | 79.68 |
| MGV Oh (2025) | 88.70 | 88.15 | 89.42 | 89.64 |
| ReMA Wan and others (2025) | 87.83 | 89.26 | 92.70 | 87.65 |
| Dataset | MetaCrit | 95% CI | Strongest baseline | ||
| TruthfulQA | 817 | 94.12 | [92.3, 95.5] | MEP 89.35 | .001 |
| CIAR | 50 | 84.00 | [71.5, 91.7] | MEP 82.00 | .79 |
| CREAK | 230 | 87.39 | [82.5, 91.1] | SOFAI-LM 84.35 | .35 |
| CrowS-Pairs | 251 | 82.07 | [76.9, 86.3] | ReMA 76.89 | .15 |
| Dataset | (p10–p90) | |||
| TruthfulQA | 204 | 0.088–0.194 | 0.46 [0.34, 0.56] | 0.37 [0.25, 0.49] |
| CREAK | 230 | 0.106–0.168 | 0.04 [ 0.16, 0.09] | 0.41 [0.29, 0.52] |
| MMLU | 98 | 0.170–0.303 | 0.04 [ 0.18, 0.25] | 0.63 [0.48, 0.75] |
| BBQ | 220 | 0.166–0.317 | 0.17 [0.04, 0.30] | 0.22 [0.08, 0.35] |
| CrowS-Pairs | 147 | 0.207–0.348 | 0.75 [0.66, 0.81] | 0.74 [0.65, 0.80] |
| Backbone | TruthfulQA | CIAR | BOLD | HONEST |
| Acc. (%) | Acc. (%) | Toxic (%) | Score | |
| MetaCrit + GPT-3.5-Turbo | 94.12 | 84 | 0.000 | 0.000 |
| MetaCrit + DeepSeek-v3 | 93.75 | 70 | 0.000 | 0.000 |
| MetaCrit + Claude-3.5-Sonnet | 97.55 | 74 | 0.000 | 0.000 |
| MetaCrit + GPT-4o | 95.83 | 96 | 0.000 | 0.000 |
| MetaCrit + GPT-4.1 | 93.83 | 96 | 0.000 | 0.000 |
| TruthfulQA | CIAR | |||
| Agreement Pattern | Acc. (%) | Acc. (%) | ||
| Both agents flag | 33 | 94.12 | 4 | 96.0 |
| Control only | 89 | 83.1 | 7 | 71.4 |
| Monitoring only | 25 | 64.0 | 0 | – |
| Neither flags | 57 | 73.7 | 1 | 0.0 |
| Failure Type | TruthfulQA | CIAR |
| Consensus collapse | 27 (56.3%) | 0 |
| Knowledge gap | 13 (27.1%) | 0 |
| Reconciliation failure | 8 (16.7%) | 8 (100%) |
| Module | Agent Content | Agent Input | Agent Output |
| Independent Agents | User Prompt Generator: | Input | Learner Profile |
| Stage Classifier: | Input | Stage | |
| Assessment: | Input | Feedback | |
| Final Response Generation: | Input , Vocab/Writing Feedback, Aggregated Prompt | Response | |
| Topic Module | Topic Identifier: | Input | Topic |
| Prompt Aggregator: | Topic, Stage Prompt | Aggregated Prompt |
| Component | Zero-Shot Single Agent | Multi-Agent | MA + MetaCrit |
| User Prompt Generator | Optional prefix | ||
| Stage Classifier | (inline) | (modular) | (modular) |
| Topic Classifier | – | ||
| Vocabulary Module | – | ||
| Writing Assessor | – | ||
| Monitoring Agent ( ) | – | – |
| Aspect | Question Numbers |
| Critical Thinking | Q3, Q4, Q9, Q14, Q16, Q19 |
| Instructiveness | Q6, Q10, Q12, Q24, Q27 |
| Interactiveness | Q11, Q17, Q18, Q22 |
| Intelligence | Q5, Q7, Q8, Q13, Q15, Q20, Q21, Q25 |