cs.AIOct 6, 2026

ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

Authors: Hanjun Luo, Xiucheng Zhang, Zhuoning Xu, Zhimu Huang, Yingbin Jin, Xinfeng Li, Hanan Salam

Organizations: New York University · New York University Abu Dhabi · The Hong Kong Polytechnic University

Abstract

As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

    May 5, 2026Jonathan Steinberg, Oren GalLanguage Model Safety EvaluationAI Coding Agents

  2. Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning

    Aug 31, 2026Fukang Zhu, Binbin Zhao, Ruixiao Lin +3Prompt SensitivityAI Coding Agents