Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance. Over 1440 cached audits from four text endpoints, the primary DeepSeek auditor reaches 100.0% closed and clean-aware code recall, but open prompts over-flag 93.8% of clean code controls, and clean-aware all-three specificity is 87.5% even where recall saturates. A clean-aware warning drops DeepSeek code false positives from 20.8% (95% CI 11.7--34.3) to 0.0% (0.0--7.4) at unchanged recall, while the budget anchor still flags 38/48 clean controls under the same prompt. Reporting recall alone would rank three of these four models identically; reporting the clean-control rate separates them by 79 points.
Figures & tables
Condition
TP/FN
FP/TN
R
FPR
Threads
open prose
39/9
18/30
81.2
37.5
37.5
open code
35/13
45/3
72.9
93.8
93.8
closed prose
48/0
1/47
100.0
2.1
2.1
closed code
48/0
10/38
100.0
20.8
20.8
clean-warn prose
48/0
0/48
100.0
0.0
0.0
clean-warn code
48/0
0/48
100.0
0.0
0.0
Table 1: Primary DeepSeek prompt-by-surface confusion matrix. Flawed cells are true positives/false negatives; clean cells are false positives/true negatives. Threads are expected unnecessary reviews per 100 clean submissions.
Model
R
FPR (95% CI)
A3
U1
U2
DeepSeek V4 Flash
100.0
0.0 (0.0–7.4)
87.5
1.00
1.00
Gemini 2.5 Flash Lite
87.5
0.0 (0.0–7.4)
85.4
0.88
0.88
GPT-4.1 mini
100.0
16.7 (8.7–29.6)
97.9
0.83
0.67
GPT-4o mini
100.0
79.2 (65.7–88.3)
75.0
0.21
−0.58
Table 2: Clean-aware code slice, four endpoints, 48 flawed and 48 clean items each. FPR intervals are 95% Wilson. A3 is all-three specificity; Uλ is Rˉ−λF .
Persona
R
FPR (95% CI)
A3
Conf.
Auditor
100.0
0.0 (0.0–7.4)
87.5
0.953
Quant reviewer
100.0
2.1 (0.4–10.9)
93.8
0.952
Skeptical reviewer
100.0
10.4 (4.5–22.2)
93.8
0.928
Risk manager
95.8
12.5 (5.9–24.7)
91.7
0.911
Table 3: Reviewer persona on the clean-aware code condition, primary model, 48 flawed and 48 clean items each. Only the assigned role changes.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Flaw class
open prose
open code
closed
Look-ahead
6/6
6/6
6/6
Survivorship
6/6
6/6
6/6
Data snooping
6/6
6/6
6/6
Overlapping-label leakage
6/6
5/6
6/6
Unrealistic fills/costs
5/6
5/6
6/6
P-hacked stop/rule
6/6
4/6
6/6
Appendix
Table 4: Primary-auditor recall by flaw class, 6 flawed items per class. Open prompts omit the taxonomy; the closed column reaches 6/6 on every class and both surfaces.