cs.AIMar 30, 2026

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

Authors: Han Wang, Yifan Sun, Brian Ko, Mann Talati, Jiawen Gong, Zimeng Li, Naicheng Yu, Xucheng Yu, +3 more

Organizations: University of Illinois Urbana-Champaign · University of Washington · University of California San Diego

Abstract

Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer faithfully reflects the actual reasons (i.e., decision-critical factors) driving the model's behavior, leading to the reduced CoT monitorability problem. This limits the use of CoTs for reliable oversight. However, a comprehensive and fully open-source benchmark for thoroughly evaluating CoT monitorability remains lacking. To address this gap, we propose MonitorBench, a systematic benchmark for evaluating CoT monitorability in LLMs. MonitorBench provides: (1) a diverse set of 1,514 test instances with carefully designed decision-critical factors across 19 tasks spanning 7 categories to characterize when CoTs can be used to monitor the factors driving LLM behavior; and (2) two prompting stress-test settings to quantify the extent to which CoT monitorability can be degraded. Extensive experiments show that CoT monitorability is a conditional property affected by the evaluated LLM, monitor LLM, and task characteristics. Across these factors, monitorability is higher when decision-critical factors shape the intermediate reasoning process, rather than merely influencing the final answer. Under stress-test prompting, most evaluated LLMs can intentionally reduce monitorability, mainly on tasks where decision-critical factors are not structurally required by the reasoning process. Overall, MonitorBench provides a basis for further research on AI control, reasoning faithfulness, stress-test monitorability, and monitoring scaffords. The code is available at https://github.com/ASTRAL-Group/MonitorBench.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?

    Aug 5, 2026Pedro Ferreira, Wilker Aziz, Ivan TitovReasoning TracesChain-of-Thought Reasoning

  2. The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

    May 27, 2026Eric Onyame, Runtao Zhou, Kowshik Thopalli +2Chain-of-Thought ReasoningSingle-Token Output Distributions

  3. Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

    Jul 9, 2026Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky +2PersuasionChain-of-Thought Reasoning