Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one CCR plus one cross-model review matches more planted errors than two CCR reviews (56.7% vs. 42.7%; Holm-adjusted p=.006), but not significantly more than two reviews by the top-tier cross-model reviewer, so model difference and reviewer capability are not separated. A lightweight cross-model reviewer scores no higher than same-model review. Withholding requirements from the reviewer raises F1 for the two lower tiers but not the top tier, in untested point estimates whose pattern depends on how failed sessions are scored. Before analysis we audited all session records, excluding one baseline run of uncertain provenance and 14 failed calls; results with all sessions are also reported. A partial check on public detector outputs from another benchmark neither replicates nor contradicts the main comparison. Records, artifacts, and scripts are available from the author on request.
Figures & tables
Condition
C
M
R
Reviewer Model
Description
SR
−
−
−
Generator (same session)
Baseline: self-review
SR2
−
−
−
Generator (2nd pass)
Repeated review (T=0)
SA
+
−
−
Generator (sub-agent)
New session with request (S varied)
CCR
+
−
+
Generator (new session)
New session, artifact only
XMR-Ge
+
+
−
Gemini 2.5 Flash
Cross-model (lightweight)
XMR-Ge-R
+
+
+
Gemini 2.5 Flash
+ information restriction
Table 1: Experimental conditions mapped to independence dimensions. Temporal (T) and structural (S) dimensions are held constant across most conditions; SR2 uniquely lacks temporal independence (repeated review), and SA uniquely varies structural independence (sub-agent role). See Sections 2.2–2.3 for detailed analysis of these dimensions. SA receives the original request with the artifact, as the unrestricted XMR conditions do, and CCR receives the artifact only, as the -R conditions do (prompts in Appendix B ). Version 1 of this paper marked CCR as unrestricted and SA as lacking context independence; both labels were inconsistent with the experiment code.
Rank
Condition
n
F1
Prec.
Rec.
pW
pt
1
XMR-GPT
86
32.3
28.6
38.4
1.00
.478
2
XMR-GPT-R
89
32.1
27.5
39.8
1.00
.524
3
XMR-GePro-R
90
29.0
29.2
29.3
1.00
.176
4
CCR
90
28.6
31.5
27.1
—
—
5
SR
60
27.1
26.2
28.3
1.00
.008
6
XMR-Ge-R
82
27.0
34.2
23.9
1.00
.232
Table 2: Performance across 10 conditions (900 sessions; n = sessions analysed after excluding SR run 3 and 14 failed cross-model sessions, Section 4.1 ; Appendix F gives the values with failures scored as zero). Values are percentages. pW : Wilcoxon signed-rank test against CCR on per-artifact means over the available runs, Holm-adjusted over nine comparisons (only SR2 differs); pt : paired t -test on run 1 only, unadjusted, as in the original CCR report.
Figure 1: Planted errors matched by CCR and by XMR-GPT-R (run 1 of each condition; 150 errors in total).
Category
CCR
GePro-R
GPT
GPT-R
Code
40.7
23.4
37.2
29.8
Document
24.5
38.1
42.1
40.7
Script
20.7
25.6
15.4
25.8
Table 3: F1 (%) by artifact category. Bold indicates best condition per category.
Error Type
CCR
XMR-Ge
GePro
GPT
FACT
48
35
48
58
CONS
37
33
38
53
CTXT
13
9
9
26
RCVR
19
7
16
34
MISS
19
17
14
21
Table 4: Detection rate (%) by error type (recall per type). Bold indicates best condition per type.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Group
n
Crit.
Major
Minor
Same-model
TP
374
22.5
32.1
45.5
FP
1,106
6.1
22.2
70.0
Cross-model
TP
771
19.7
50.3
30.0
FP
1,988
12.9
41.8
45.3
Appendix
Table 5: Severity labels attached by reviewers (% of findings). Analysed sessions as in Table 2 , minus 26 sessions (24 SR, 2 XMR-Ge-R) whose findings were recovered from raw output and cannot be aligned with their labels. Because 24 of the 26 are SR sessions, SR contributes 36 of its 60 analysed sessions to the same-model group, and the cross-model false positives number 1,988 rather than the 1,995 of Section 5.8 . Twenty same-model false positives carry other labels and are omitted from the percentages’ numerators; sums may differ by 0.1 from rounding. Descriptive; not tested.
Figure 2: F1 scores for all 10 conditions (ranked). The asterisk marks the only condition significantly different from CCR (SR2; Wilcoxon signed-rank on per-artifact means over the available runs, Holm-adjusted p=0.011 ); full statistics in Table 2 .
Condition
All sessions
Excluded
n excl.
XMR-GPT
30.9
32.3
86
XMR-GPT-R
31.8
32.1
89
XMR-GePro-R
29.0
29.0
90
CCR
28.6
28.6
90
SR
24.5
27.1
60
XMR-Ge-R
24.6
27.0
82
Appendix
Table 6: F1 (%) with all 900 sessions and after the exclusions. With all sessions, the primary test against CCR (Holm over nine comparisons) finds only SR2 different ( p=0.011 ); SR’s unadjusted p is 0.021 (0.26 after excluding its run 3), and XMR-GPT vs. XMR-Ge has unadjusted p=0.071 (0.020 after excluding failures).
Comparison axis
Mean Jaccard
Range
Inter-perspective (same model)
0.228
0.143–0.331
Inter-model (same perspective)
0.152
0.057–0.259
Appendix
Table 7: Perspective vs. model diversification Jaccard overlap (lower means less overlap).