The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability
Abstract
Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hyperproperties, witnessed only by two executions. The standard consequence is a binary impossibility: one trace cannot decide them. We replace the binary with a measurement. A tight bound puts the balanced accuracy of any single-trace monitor at , turning undecidability into a graded detectability frontier and defining an oversight gap: a monitor's shortfall below it. On a leak family with closed-form , nine LLM monitors are optimal at but capture little signal as grows; at , where a 20-line membership check scores , they average . That shortfall is mostly not capability: naming what to check closes of it while leaving the control at chance. The same split runs through a factorial: an imagined second run leaves monitors at chance () while the same rule on an executed second run reaches , and a stored oracle without a comparison procedure yields only . Information and procedure are each necessary and neither is capability. Under nondeterminism, replay tracks a closed-form -replay curve only under the right projection, and a projection frontier shows the resulting dilemma is forced: narrow misses of off-channel leaks, broad flags of clean traffic, and attainable accuracy decays like in the benign-variation rate and the channel count. Finally, two frontier LLM judges certified an earlier version of our own benchmark as sound while a sign test found a directional bias () that invalidated three of our findings. Construction validity for hyperproperty benchmarks should be proved mechanically, not audited by models.