ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
Organizations: New York University · New York University Abu Dhabi · The Hong Kong Polytechnic University
Abstract
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.
Figures & tables
| Treatment | Evidence entails | #Situations | Representative situations |
| Avoidance | 8 | Reconstructible artifacts; serialized access | |
| Transfer | 8 | Continuous integration; human review | |
| Mitigation | 11 | Stated count or batch size; requested verification tier | |
| Acceptance | 17 | Final identity or status; deletion |
| TSR (%) | VR (%) | ER (%) | |||||||||||
| Model | All | All | Avo | Tra | Mit | Acc | All | Avo | Tra | Mit | Acc | ||
| Opus-5 | 95.8 | 96.8 | 94.7 | 58.7 | 48.7 | 59.3 | 56.0 | 70.7 | 25.4 | 25.6 | -3.5 | 41.7 | 26.4 |
| Sonnet-5 | 93.4 | 95.5 | 91.3 | 22.2 | 23.3 | 10.7 | 35.3 | 19.3 | 51.3 | 27.2 | 23.3 | 64.2 | 48.3 |
| Sonnet-4.6 | 90.3 | 91.3 | 89.3 | 11.2 | 12.7 | 3.3 | 15.3 | 13.3 | 62.6 | 13.7 | 38.9 | 74.4 | 66.7 |
| Haiku-4.5 | 88.3 | 89.7 | 87.0 | 17.7 | 28.7 | 12.0 | 24.7 | 5.3 | 65.5 | 0.0 | 28.2 | 65.7 | 93.9 |
| Claude | 92.0 | 93.3 | 90.6 | 27.4 | 28.3 | 21.3 | 32.8 | 27.2 | 51.2 | 16.6 | 21.7 | 61.5 | 58.8 |
| Excess rate (%) | (pp) | ER (%) | |||||
| Model | Est. | 95% CI | Est. | 95% CI | |||
| Opus-5 | 58.7 | 78.7 | 20.0 | 25.4 | |||
| Sonnet-5 | 22.2 | 45.5 | 23.3 | 51.3 | |||
| Sonnet-4.6 | 11.2 | 30.0 | 18.8 | 62.8 | |||
| Haiku-4.5 | 17.7 | 51.3 | 33.7 | 65.6 | |||
| 5.6-Sol | 21.0 | 52.5 | 31.5 | 60.0 | |||
| #Records | Sat. | Accept (%) | |||||||
| Group | w/ | w/o | w/ | w/o | 95% CI | w/ | w/o | ||
| All | 400 | 400 | 2.58 | 3.85 | 1.27 | 38% | 78% | ||
| Avoidance | 100 | 100 | 2.72 | 3.82 | 1.10 | 45% | 78% | ||
| Transfer | 100 | 100 | 2.64 | 3.84 | 1.20 | 40% | 77% | ||
| Mitigation | 100 | 100 | 2.35 | 3.90 | 1.55 | 28% | 80% | ||
| Acceptance | 100 | 100 | 2.61 | 3.84 | 1.23 | 39% | 77% | ||
| Treatment | (VR) | ER | ||
| Avoidance | 39.2 | 30.2 9.0 | 21.3 | |
| Transfer | 21.8 | 20.9 0.9 | 7.9 | |
| Mitigation | 78.5 | 26.4 52.1 | 67.4 | |
| Acceptance | 54.0 | 19.3 34.7 | 65.4 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Recurring situation | Typical unnecessary treatment | Tasks |
| Reconstructible artifacts | Backups or integrity checks for artifacts regenerated from versioned inputs. | 6 |
| Producer-controlled formats | Schema walls, broad exception handling, or silent defaults around controlled outputs. | 7 |
| Serialized access | Locks, atomic writes, or conflict retries on a single-writer path. | 5 |
| Internal-only inputs | Null, type, or range guards for values covered by an internal invariant. | 11 |
| One-off or short-lived work | Checkpoints, retries, or configuration layers for disposable work. | 7 |
| Local moves and renames | Backups or before–after hashes within the same tracked workspace. | 4 |
| Recurring situation | Typical unnecessary treatment | Tasks |
| Existing supervisor or health monitor | Another poller, monitor, or alert for already monitored work. | 5 |
| Continuous integration | Another workflow, gate, runner, or full-suite run for checks assigned to CI. | 11 |
| Validation manifest | Recomputing entries or creating a second maintained manifest. | 4 |
| Prior validation record | Rechecking completed items instead of addressing the recorded gap. | 7 |
| Upstream or receiver validation | Another checksum or handoff protocol around an existing validator. | 8 |
| Version-control recovery | Backup files, directories, stubs, or commented copies of tracked code. | 6 |
| Recurring situation | Typical unnecessary treatment | Tasks |
| Stated frequency or interval | Polling or reporting more frequently than requested. | 2 |
| Stated count, concurrency, or batch size | Changing the stated number of workers, attempts, or items. | 6 |
| Named files, modules, or functions | Editing or validating outside the named scope. | 6 |
| Requested length or form | Expanding a bounded deliverable into a substantially larger artifact. | 6 |
| Requested verification tier | Expanding a targeted check into broader or repeated verification. | 15 |
| Stated time or lifetime | Turning a bounded trial into a persistent process or full-scale run. | 3 |
| Recurring situation | Typical unnecessary treatment | Tasks |
| Final identity or status | Relabeling an official result as provisional or awaiting approval. | 6 |
| Chosen disposition | Preserving or rerunning an artifact after another disposition was selected. | 3 |
| Selected solution or parameter | Replacing the selected choice with a conservative substitute. | 2 |
| Declared scope or prohibition | Broadening a narrow exclusion into additional restrictions. | 3 |
| User-owned next step | Performing review, publication, tagging, or comprehensive testing retained by the user. | 2 |
| Revised instruction | Continuing to follow a restriction the user has superseded. | 2 |
| Score | Satisfaction | Participant-facing interpretation |
| 1 | Very dissatisfied | The code would be hard to maintain, the run took far longer than the task required, or the output would need extensive corrections before use. |
| 2 | Dissatisfied | The run adds noticeable maintenance burden, wasted time, or several corrections before the output can be used. |
| 3 | Neutral | The run is usable, but its extra code, time, or required corrections offset part of its value. |
| 4 | Satisfied | The run needs at most a minor correction and adds little maintenance burden or wasted time. |
| 5 | Very satisfied | The code is easy to maintain, the runtime fits the task, and the output can be used without corrections. |
| Model | Context window | Maximum output |
| Claude Opus 5, Claude Sonnet 5 | 1M | 64K |
| Claude Sonnet 4.6, Claude Haiku 4.5 | 200K | 32K |
| All evaluated GPT models | 400K (272K usable input) | 128K |
| Setting | Value |
| Evaluator model | GPT-5.6 Sol |
| Retrieved exemplars | 4 per trajectory |
| Maximum contrast pairs | 2 |
| Retrieval encoder | nomic-embed-text-v1.5 ( Nussbaum et al., 2024 ) |
| Embedding dimension | 512 |
| Judge | Accuracy | Precision | Recall | F1 | |
| GPT-5.6 Sol | 0.965 | 0.93 | 0.971 | 0.962 | 0.966 |
| Claude Sonnet 5 | 0.945 | 0.89 | 0.943 | 0.952 | 0.947 |
| Agent configuration | Primary VR | Audit VR | Centered difference (95% CI) |
| Claude Opus 5 | 58.0% | 60.5% | pp |
| Claude Sonnet 5 | 24.0% | 25.8% | pp |
| Claude Sonnet 4.6 | 14.0% | 14.5% | pp |
| Claude Haiku 4.5 | 18.0% | 17.8% | pp |
| GPT-5.6 Sol | 23.0% | 24.3% | pp |
| GPT-5.6 Terra | 17.0% | 17.0% | pp |
| TSR (%) | VR (%) | ER (%) | |||||||||||
| Model | All | All | Avo | Tra | Mit | Acc | All | Avo | Tra | Mit | Acc | ||
| Opus-5 | 95.0 | 96.0 | 94.0 | 61.0 | 54.0 | 58.0 | 58.0 | 74.0 | 25.6 | 25.0 | -3.6 | 42.0 | 26.0 |
| Sonnet-5 | 92.0 | 93.5 | 90.5 | 22.0 | 22.0 | 12.0 | 34.0 | 20.0 | 51.6 | 26.7 | 25.0 | 66.0 | 44.4 |
| Sonnet-4.6 | 90.5 | 92.5 | 88.5 | 12.0 | 14.0 | 4.0 | 16.0 | 14.0 | 62.5 | 12.5 | 33.3 | 75.0 | 66.7 |
| Haiku-4.5 | 88.8 | 90.0 | 87.5 | 17.5 | 30.0 | 10.0 | 24.0 | 6.0 | 67.3 | 0.0 | 28.6 | 65.7 | 94.0 |
| Claude | 91.6 | 93.0 | 90.1 | 28.1 | 30.0 | 21.0 | 33.0 | 28.5 | 51.8 | 16.1 | 20.8 | 62.2 | 57.8 |
| TSR (%) | VR (%) | ER (%) | |||||||||||
| Model | All | All | Avo | Tra | Mit | Acc | All | Avo | Tra | Mit | Acc | ||
| Opus-5 | 95.5 | 95.5 | 95.5 | 56.0 | 42.0 | 58.0 | 52.0 | 72.0 | 25.8 | 27.6 | -3.6 | 42.2 | 26.5 |
| Sonnet-5 | 92.8 | 96.0 | 89.5 | 22.5 | 28.0 | 12.0 | 34.0 | 16.0 | 51.6 | 26.3 | 25.0 | 65.3 | 52.9 |
| Sonnet-4.6 | 89.2 | 91.0 | 87.5 | 10.0 | 12.0 | 4.0 | 12.0 | 12.0 | 60.8 | 14.3 | 33.3 | 73.9 | 66.7 |
| Haiku-4.5 | 88.0 | 89.5 | 86.5 | 18.0 | 32.0 | 10.0 | 24.0 | 6.0 | 66.4 | 0.0 | 28.6 | 65.7 | 93.9 |
| Claude | 91.4 | 93.0 | 89.8 | 26.6 | 28.5 | 21.0 | 30.5 | 26.5 | 51.2 | 17.1 | 20.8 | 61.8 | 60.0 |
| TSR (%) | VR (%) | ER (%) | |||||||||||
| Model | All | All | Avo | Tra | Mit | Acc | All | Avo | Tra | Mit | Acc | ||
| Opus-5 | 96.8 | 99.0 | 94.5 | 59.0 | 50.0 | 62.0 | 58.0 | 66.0 | 24.8 | 24.2 | -3.3 | 40.8 | 26.7 |
| Sonnet-5 | 95.5 | 97.0 | 94.0 | 22.0 | 20.0 | 8.0 | 38.0 | 22.0 | 50.6 | 28.6 | 20.0 | 61.2 | 47.6 |
| Sonnet-4.6 | 91.2 | 90.5 | 92.0 | 11.5 | 12.0 | 2.0 | 18.0 | 14.0 | 64.6 | 14.3 | 50.0 | 74.3 | 66.7 |
| Haiku-4.5 | 88.2 | 89.5 | 87.0 | 17.5 | 24.0 | 16.0 | 26.0 | 4.0 | 62.8 | 0.0 | 27.3 | 65.8 | 93.9 |
| Claude | 92.9 | 94.0 | 91.9 | 27.5 | 26.5 | 22.0 | 35.0 | 26.5 | 50.7 | 16.8 | 23.5 | 60.5 | 58.7 |
| Agent configuration | Run 1 | Run 2 | Run 3 | Pooled | Max–min |
| Claude Opus 5 | 21.0 pp | 19.5 pp | 19.5 pp | 20.0 pp | 1.5 pp |
| Claude Sonnet 5 | 23.5 pp | 24.0 pp | 22.5 pp | 23.3 pp | 1.5 pp |
| Claude Sonnet 4.6 | 20.0 pp | 15.5 pp | 21.0 pp | 18.8 pp | 5.5 pp |
| Claude Haiku 4.5 | 36.0 pp | 35.5 pp | 29.5 pp | 33.7 pp | 6.5 pp |
| GPT-5.6 Sol | 34.5 pp | 34.0 pp | 26.0 pp | 31.5 pp | 8.5 pp |
| GPT-5.6 Terra | 21.5 pp | 24.0 pp | 28.0 pp | 24.5 pp | 6.5 pp |
| Model | Total input | Output | Runtime (h) |
| Claude Opus 5 | 306.3M | 3.6M | 30.9 |
| Claude Sonnet 5 | 391.8M | 2.4M | 16.5 |
| Claude Sonnet 4.6 | 278.7M | 2.1M | 15.9 |
| Claude Haiku 4.5 | 357.9M | 3.3M | 13.5 |
| GPT-5.6 Sol | 431.1M | 2.4M | 70.2 |
| GPT-5.6 Terra | 187.2M | 1.8M | 18.9 |