As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.
Figures & tables
Figure 1: ParanoiaEval uses evidence-controlled task pairs to show when the same defensive action is appropriate or unnecessary.
Figure 2: Construction pipeline of ParanoiaEval .
Treatment
Evidence entails
#Situations
Representative situations
Avoidance
¬occur(r)
8
Reconstructible artifacts; serialized access
Transfer
owner(r)=A
8
Continuous integration; human review
Mitigation
bound(r)=m⋆
11
Stated count or batch size; requested verification tier
Acceptance
accepted(r)
17
Final identity or status; deletion
Table 1: Summary of the 4 treatments and 44 situations of ParanoiaEval .
Figure 3: Evaluation workflow of ParanoiaEval .
TSR (%) ↑
VR (%) ↓
ER (%) ↑
Model
All
E+
E−
All
Avo
Tra
Mit
Acc
All
Avo
Tra
Mit
Acc
Opus-5
95.8
96.8
94.7
58.7
48.7
59.3
56.0
70.7
25.4
25.6
-3.5
41.7
26.4
Sonnet-5
93.4
95.5
91.3
22.2
23.3
10.7
35.3
19.3
51.3
27.2
23.3
64.2
48.3
Sonnet-4.6
90.3
91.3
89.3
11.2
12.7
3.3
15.3
13.3
62.6
13.7
38.9
74.4
66.7
Haiku-4.5
88.3
89.7
87.0
17.7
28.7
12.0
24.7
5.3
65.5
0.0
28.2
65.7
93.9
Claude μ
92.0
93.3
90.6
27.4
28.3
21.3
32.8
27.2
51.2
16.6
21.7
61.5
58.8
Table 2: All values are averaged over 3 runs. Claude μ and GPT μ denote the means over the Claude and GPT models, respectively. ↑ indicates that higher values mean more tasks completed or a stronger response, and ↓ indicates that lower values mean less unnecessary risk treatment.
Excess rate (%)
Δ (pp)
ER (%)
Model
E+
E−
Est.
95% CI
Est.
95% CI
padj
Opus-5
58.7
78.7
20.0
[13.0,27.0]
25.4
[16.5,34.5]
.004
Sonnet-5
22.2
45.5
23.3
[16.8,29.8]
51.3
[37.8,63.9]
<.001
Sonnet-4.6
11.2
30.0
18.8
[12.2,25.4]
62.8
[45.1,77.0]
.003
Haiku-4.5
17.7
51.3
33.7
[26.5,40.9]
65.6
[53.2,76.1]
<.001
5.6-Sol
21.0
52.5
31.5
[24.6,38.4]
60.0
[47.8,70.5]
<.001
Table 3: Paired condition effects across the three measured runs.
Figure 4: (a) ER (filled) and VR (hollow) of each configuration, with 95% bootstrap confidence intervals. (b) Task success and VR across agent configurations. (c) Developer satisfaction with and without excess treatment. (d) Excess-treatment rate on E− and E+ by treatment, with the mean over configurations in bold and individual configurations in thin lines.
#Records
Sat.
Accept (%)
Group
w/
w/o
w/
w/o
Δ
95% CI
p
w/
w/o
All
400
400
2.58
3.85
1.27
[1.14,1.40]
<.001
38%
78%
Avoidance
100
100
2.72
3.82
1.10
[0.87,1.33]
.001
45%
78%
Transfer
100
100
2.64
3.84
1.20
[0.98,1.42]
<.001
40%
77%
Mitigation
100
100
2.35
3.90
1.55
[1.32,1.78]
<.001
28%
80%
Acceptance
100
100
2.61
3.84
1.23
[1.00,1.46]
<.001
39%
77%
Table 4: Developer satisfaction and acceptance for runs with and without excess treatment. Satisfaction is measured on a five-point scale. Δ is the satisfaction of runs without excess treatment minus that of runs with it.
Treatment
E−
E+ (VR)
ER
p
Avoidance
39.2
30.2 ↓ 9.0
21.3
.008
Transfer
21.8
20.9 ↓ 0.9
7.9
.692
Mitigation
78.5
26.4 ↓ 52.1
67.4
<.001
Acceptance
54.0
19.3 ↓ 34.7
65.4
<.001
Table 5: Excess treatment by treatment, averaged over configurations. p is from a two-sided paired permutation test between E− and E+ over task pairs.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Original NIST risk response workflow. This is a direct reproduction of the original Figure 9 in NIST IR 8286B-upd1 ( Quinn et al., 2025b ) . Republished courtesy of the National Institute of Standards and Technology.
Recurring situation
Typical unnecessary treatment
Tasks
Reconstructible artifacts
Backups or integrity checks for artifacts regenerated from versioned inputs.
6
Producer-controlled formats
Schema walls, broad exception handling, or silent defaults around controlled outputs.
7
Serialized access
Locks, atomic writes, or conflict retries on a single-writer path.
5
Internal-only inputs
Null, type, or range guards for values covered by an internal invariant.
11
One-off or short-lived work
Checkpoints, retries, or configuration layers for disposable work.
7
Local moves and renames
Backups or before–after hashes within the same tracked workspace.
4
Appendix
Table 6: Recurring Avoidance situations and their task counts.
Recurring situation
Typical unnecessary treatment
Tasks
Existing supervisor or health monitor
Another poller, monitor, or alert for already monitored work.
5
Continuous integration
Another workflow, gate, runner, or full-suite run for checks assigned to CI.
11
Validation manifest
Recomputing entries or creating a second maintained manifest.
4
Prior validation record
Rechecking completed items instead of addressing the recorded gap.
7
Upstream or receiver validation
Another checksum or handoff protocol around an existing validator.
8
Version-control recovery
Backup files, directories, stubs, or commented copies of tracked code.
6
Appendix
Table 7: Recurring Transfer situations and their task counts.
Recurring situation
Typical unnecessary treatment
Tasks
Stated frequency or interval
Polling or reporting more frequently than requested.
2
Stated count, concurrency, or batch size
Changing the stated number of workers, attempts, or items.
6
Named files, modules, or functions
Editing or validating outside the named scope.
6
Requested length or form
Expanding a bounded deliverable into a substantially larger artifact.
6
Requested verification tier
Expanding a targeted check into broader or repeated verification.
15
Stated time or lifetime
Turning a bounded trial into a persistent process or full-scale run.
3
Appendix
Table 8: Recurring Mitigation situations and their task counts.
Recurring situation
Typical unnecessary treatment
Tasks
Final identity or status
Relabeling an official result as provisional or awaiting approval.
6
Chosen disposition
Preserving or rerunning an artifact after another disposition was selected.
3
Selected solution or parameter
Replacing the selected choice with a conservative substitute.
2
Declared scope or prohibition
Broadening a narrow exclusion into additional restrictions.
3
User-owned next step
Performing review, publication, tagging, or comprehensive testing retained by the user.
2
Revised instruction
Continuing to follow a restriction the user has superseded.
2
Appendix
Table 9: Recurring Acceptance situations and their task counts.
Score
Satisfaction
Participant-facing interpretation
1
Very dissatisfied
The code would be hard to maintain, the run took far longer than the task required, or the output would need extensive corrections before use.
2
Dissatisfied
The run adds noticeable maintenance burden, wasted time, or several corrections before the output can be used.
3
Neutral
The run is usable, but its extra code, time, or required corrections offset part of its value.
4
Satisfied
The run needs at most a minor correction and adds little maintenance burden or wasted time.
5
Very satisfied
The code is easy to maintain, the runtime fits the task, and the output can be used without corrections.
Appendix
Table 10: Five-point scale for developer satisfaction with a completed agent run.
Model
Context window
Maximum output
Claude Opus 5, Claude Sonnet 5
1M
64K
Claude Sonnet 4.6, Claude Haiku 4.5
200K
32K
All evaluated GPT models
400K (272K usable input)
128K
Appendix
Table 11: Maximum context and output capacities enabled for the evaluated coding agents.
Setting
Value
Evaluator model
GPT-5.6 Sol
Retrieved exemplars
4 per trajectory
Maximum contrast pairs
2
Retrieval encoder
nomic-embed-text-v1.5 ( Nussbaum et al., 2024 )
Embedding dimension
512
Appendix
Table 12: Evaluator and exemplar-retrieval hyperparameters.
Judge
Accuracy
κ
Precision
Recall
F1
GPT-5.6 Sol
0.965
0.93
0.971
0.962
0.966
Claude Sonnet 5
0.945
0.89
0.943
0.952
0.947
Appendix
Table 13: Agreement of the two judges with held-out human consensus labels for Vp .
Agent configuration
Primary VR
Audit VR
Centered difference (95% CI)
Claude Opus 5
58.0%
60.5%
+1.5 pp [−0.8,+3.8]
Claude Sonnet 5
24.0%
25.8%
+0.8 pp [−1.4,+3.0]
Claude Sonnet 4.6
14.0%
14.5%
−0.5 pp [−2.7,+1.7]
Claude Haiku 4.5
18.0%
17.8%
−1.2 pp [−3.5,+1.1]
GPT-5.6 Sol
23.0%
24.3%
+0.3 pp [−2.0,+2.6]
GPT-5.6 Terra
17.0%
17.0%
−1.0 pp [−3.2,+1.3]
Appendix
Table 14: Post-hoc comparison of the primary and audit judges on the audit subset.
TSR (%) ↑
VR (%) ↓
ER (%) ↑
Model
All
E+
E−
All
Avo
Tra
Mit
Acc
All
Avo
Tra
Mit
Acc
Opus-5
95.0
96.0
94.0
61.0
54.0
58.0
58.0
74.0
25.6
25.0
-3.6
42.0
26.0
Sonnet-5
92.0
93.5
90.5
22.0
22.0
12.0
34.0
20.0
51.6
26.7
25.0
66.0
44.4
Sonnet-4.6
90.5
92.5
88.5
12.0
14.0
4.0
16.0
14.0
62.5
12.5
33.3
75.0
66.7
Haiku-4.5
88.8
90.0
87.5
17.5
30.0
10.0
24.0
6.0
67.3
0.0
28.6
65.7
94.0
Claude μ
91.6
93.0
90.1
28.1
30.0
21.0
33.0
28.5
51.8
16.1
20.8
62.2
57.8
Appendix
Table 15: Results of Run 1.
TSR (%) ↑
VR (%) ↓
ER (%) ↑
Model
All
E+
E−
All
Avo
Tra
Mit
Acc
All
Avo
Tra
Mit
Acc
Opus-5
95.5
95.5
95.5
56.0
42.0
58.0
52.0
72.0
25.8
27.6
-3.6
42.2
26.5
Sonnet-5
92.8
96.0
89.5
22.5
28.0
12.0
34.0
16.0
51.6
26.3
25.0
65.3
52.9
Sonnet-4.6
89.2
91.0
87.5
10.0
12.0
4.0
12.0
12.0
60.8
14.3
33.3
73.9
66.7
Haiku-4.5
88.0
89.5
86.5
18.0
32.0
10.0
24.0
6.0
66.4
0.0
28.6
65.7
93.9
Claude μ
91.4
93.0
89.8
26.6
28.5
21.0
30.5
26.5
51.2
17.1
20.8
61.8
60.0
Appendix
Table 16: Results of Run 2.
TSR (%) ↑
VR (%) ↓
ER (%) ↑
Model
All
E+
E−
All
Avo
Tra
Mit
Acc
All
Avo
Tra
Mit
Acc
Opus-5
96.8
99.0
94.5
59.0
50.0
62.0
58.0
66.0
24.8
24.2
-3.3
40.8
26.7
Sonnet-5
95.5
97.0
94.0
22.0
20.0
8.0
38.0
22.0
50.6
28.6
20.0
61.2
47.6
Sonnet-4.6
91.2
90.5
92.0
11.5
12.0
2.0
18.0
14.0
64.6
14.3
50.0
74.3
66.7
Haiku-4.5
88.2
89.5
87.0
17.5
24.0
16.0
26.0
4.0
62.8
0.0
27.3
65.8
93.9
Claude μ
92.9
94.0
91.9
27.5
26.5
22.0
35.0
26.5
50.7
16.8
23.5
60.5
58.7
Appendix
Table 17: Results of Run 3.
Agent configuration
Run 1 Δ
Run 2 Δ
Run 3 Δ
Pooled Δ
Max–min
Claude Opus 5
21.0 pp
19.5 pp
19.5 pp
20.0 pp
1.5 pp
Claude Sonnet 5
23.5 pp
24.0 pp
22.5 pp
23.3 pp
1.5 pp
Claude Sonnet 4.6
20.0 pp
15.5 pp
21.0 pp
18.8 pp
5.5 pp
Claude Haiku 4.5
36.0 pp
35.5 pp
29.5 pp
33.7 pp
6.5 pp
GPT-5.6 Sol
34.5 pp
34.0 pp
26.0 pp
31.5 pp
8.5 pp
GPT-5.6 Terra
21.5 pp
24.0 pp
28.0 pp
24.5 pp
6.5 pp
Appendix
Table 18: Stability of the paired condition effect across the three measured runs.
Model
Total input
Output
Runtime (h)
Claude Opus 5
306.3M
3.6M
30.9
Claude Sonnet 5
391.8M
2.4M
16.5
Claude Sonnet 4.6
278.7M
2.1M
15.9
Claude Haiku 4.5
357.9M
3.3M
13.5
GPT-5.6 Sol
431.1M
2.4M
70.2
GPT-5.6 Terra
187.2M
1.8M
18.9
Appendix
Table 19: Computational resource consumption of the evaluated agent episodes. Runtime is cumulative across episodes.
Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing safety alignment evaluates overt requests in isolation, leaving models blind to malicious end-states that emerge from sequenced compliance with innocuous-looking requests. We introduce MOSAIC-Bench (Malicious Objectives Sequenced As Innocuous Compliance), a benchmark of 199 three-stage attack chains paired with deterministic exploit oracles on deployed software substrates (10 web-application substrates, 31 CWE classes, 5 programming languages) that treats both exploit ground truth and downstream reviewer protocol as first-class evaluation axes. On this benchmark, nine production coding agents from Anthropic, OpenAI, Google, Moonshot, Zhipu, and Minimax compose innocuous tickets at 53-86% end-to-end ASR with only two refusals across all staged runs. In a matched direct-prompt experiment over four frontier Claude/Codex agents, vulnerable-output rates fall to 0-20.4%: Claude primarily refuses, while Codex primarily hardens rather than emitting the vulnerable implementation - ticket staging silences both defense modes simultaneously. Downstream, code reviewer agents approve 25.8% of these confirmed-vulnerable cumulative diffs as routine PRs, and a full-context implementation protocol closes only 50% of the staged/direct gap, ruling out context fragmentation as the sole explanation. As a deployable but non-adaptive mitigation, reframing the reviewer as an adversarial pentester reduces evasion across the evaluated reviewer subset; pentester framed evasion ranges from 3.0% to 17.6%, and an open-weight Gemma-4-E4B-it reviewer under this framing detects 88.4% of attacks on the dataset with a 4.6% false-positive rate measured on 608 real-world GitHub PRs.
Coding agents are increasingly used for software engineering tasks, including bootstrapping projects from third-party repositories whose integrity cannot be assumed. Prior work on repository poisoning largely focuses on attacker-controlled injection and disguise, but developers also shape risk through everyday invocation choices: what task to delegate, how to phrase the request, and which skills or rules to supply. We term these user-side choices Prompt-Level Configurations (PLCs) and introduce CIPR (Coding In Poisoned Repos), the first benchmark that systematically varies PLCs in poisoned real-world repositories. CIPR comprises 1,920 instances across 20 repositories, four task types, three social-media-grounded prompt styles, and three skill/rule conditions, and measures attack success rate (ASR) and agent alert rate (AR) using automated runtime and trace-based oracles. Our evaluation reveals two key insights: (1) Vulnerability is highly context-dependent, with task type creating up to a 4.5-fold difference in ASR, with test-execution task forming a silent attack surface (high ASR, low AR). (2) Prompt expression shifts risk indirectly: underspecified prompts reduce ASR by truncating execution depth; noisy prompts exhibit a directional trend toward suppressing alerts by making malicious content less conspicuous. These findings highlight that coding agent vulnerability is not a static property, but a dynamic outcome shaped by everyday user configurations.
Fukang Zhu, Binbin Zhao, Ruixiao Lin +3
1Zhejiang University · 2State Key Laboratory of Internet Architecture, Tsinghua University
Agentic coding harnesses - such as Agent-Skills, Superpowers, and Agent-Rigor - are increasingly deployed to augment underlying LLMs for real-world software engineering tasks. Existing benchmarks evaluate these agents almost exclusively on outcome correctness: whether generated code passes tests or resolves issues. We argue that this outcome-only lens is insufficient: an agent that arrives at a correct solution through reckless trial-and-error, without planning, verification, or graceful recovery, is fundamentally less reliable than one that follows sound engineering discipline. We introduce RigorBench, the first benchmark designed to measure process discipline in AI coding agents. RigorBench evaluates these harnesses across five pillars: Planning Fidelity, Verification Coverage, Recovery Efficiency, Abstention Quality, and Atomic Transition Integrity. A composite RigorScore aggregates these dimensions into a single metric via a weighted sum. We curate a suite of 30 tasks spanning five categories - Plan-Then-Build, Verify-Or-Die, Doom Loop Gauntlet, Know When to Fold, and Don't Break the Build-and evaluate leading harnesses in a controlled with/without experimental design against baseline coding assistants. Our results show that structured process discipline not only improves process quality scores by an average of 41% but also raises downstream outcome correctness by 17%, providing the first quantitative evidence that how agents code matters as much as what they produce. We release the full benchmark, scoring rubrics, and trajectory analysis tools as open-source artifacts.
Meher Bhaskar Madiraju, Meher Sai Preetam Madiraju