cs.AIJun 9, 2026

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Authors: Joachim SchaefferThomas JiralerspongAlexander PanfilovGuillaume LajoieJonas GeipingYoshua BengioRoland S. Zimmermann

Organizations: 1MATS · 2Mila – Quebec AI Institute · 3Universit´e de Montr´eal · 4Astra Fellowship · 5ELLIS Institute Tübingen, MPI for Intelligent Systems & Tübingen AI Center · 6LawZero · 7Google DeepMind

Abstract

AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the trusted model detects such an intervention, it may infer properties of the monitor and adapt to evade control. We introduce \textbf{CIAware-Bench}, a benchmark for measuring \textbf{c}ontrol \textbf{i}ntervention (CI) awareness across frontier models. CIAware-Bench tests whether models can distinguish their own trajectories from those modified by a control intervention. The benchmark is comprised of a suite of four task domains (essay writing, BigCodeBench, Bash Arena, and SHADE-Arena), while varying trajectory watermarking, side-task presence, and the control protocol. Evaluating eleven frontier models, we find low to moderate CI awareness under default settings (up to 0.87; random chance balanced binary classification accuracy is 0.5) with substantial variation across task domains and model pairs. Detection is generally easier across model families, suggesting that models exploit provider-specific differences in style or post-training. Overall, CI awareness is not a fixed model-level property, and should be measured for each new model release and deployment scenario. We release CIAware-Bench to track CI awareness and inform control protocols whose interventions are harder to detect.

Explore similar work

May 21, 2026cs.LG

Decomposing and Measuring Evaluation Awareness

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a shared foundation, conflating properties of the evaluation with properties of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component (how recognizable the task is) and a model component that separates recognition from propensity to act on it. We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring. Across nine frontier models and four benchmarks, recognition rates depend on the specific pairing of model and benchmark rather than on either in isolation. Recognition rarely leads to behavioral change, and when it does, the direction depends on the type of evaluation perceived. Models are also more sensitive to safety than capability evaluations, placing safety benchmark validity at greater risk. To study which factors each model is sensitive to and how they interact, we propose \textbf{EvalAwareBench}, a factor-controlled benchmark of 100 paired safety-capability tasks where each of the eight factors can be independently toggled, varying evaluative signals while holding the underlying request fixed. Through EvalAwareBench, we find that no single factor uniformly affects all models, but stacking factors progressively raises evaluation awareness across all of them. Our framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, pointing to behavioral consistency under recognition as a promising path forward.
Changling Li, Terry Jingchen Zhang, Jie Zhang +3
Jul 30, 2026cs.AI

InfoOps Bench: A live information operations safety benchmark

In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for use by authoritarian state "information operations": intentional, coordinated activities by one state to influence public opinion and information ecosystems in another state. These information operations are a well documented, persistent threat against contemporary democracy. Our benchmark is based on real examples from over 2,100 information operations drawn from a live monitoring pipeline which tracks online information assets with links to authoritarian regimes. Alongside this paper, we also release a companion website that updates the benchmark weekly with new claims. The dynamic nature of this public facing benchmark makes it resistant to saturation. In the benchmark, we test 17 models from 8 providers across four prompt framings. We find that most models can be co-opted for information operations at least some of the time. Integrity scores, defined as the share of judged responses in which the model neither preserved nor amplified the claim, range from 9.3% to 91%, an 81.7-percentage-point spread not explained by model size. Models approach participation in information operations in a variety of ways. Some models fabricate details and produce output more harmful than the original input claim; others make claims less harmful even while complying and producing some output. Fact-checking rates vary from 3.2% to 80.8%. Integrity against information operations is at least partly related to refusal to produce content even for benign claims, illustrating the challenge of balancing model usability with safety. Overall, our results show the potential for contemporary information operations to be substantially aided by frontier AI models.
Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright +1
Jun 10, 2026cs.LG

Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents

Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and untrusted models may render trusted models unreliable monitors. We introduce \emph{bootstrapped monitoring}, a protocol that addresses this by inserting a stronger, intermediate untrusted model with transparent chain-of-thought reasoning into the oversight chain. The untrusted monitor (UmU_m) evaluates the agent's actions, while a weaker trusted model (TT) oversees UmU_m's reasoning to detect collusion. We evaluate bootstrapped monitoring on multi-turn software engineering tasks (BashArena) across multiple agents and monitors. Bootstrapped monitoring substantially improves catch rates over trusted-only monitoring, even when the untrusted monitor actively colludes with the agent, provided we have access to its raw chain-of-thought. Our results suggest that bootstrapped monitoring can extend the useful lifetime of trusted models in control as AI capabilities advance.
Frank Xiao, Mary Phuong