The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Experiments on representative LLM-based reviewers reveal two limitations of static evaluation: stratified vulnerability, where perturbation effects depend on whether the paper's original review score is high or low, and perturbation undercoverage, where a single template misses vulnerabilities exposed by diverse realizations. To address these limitations, we propose SCOPE-Fuzzer, a strategy-aware fuzzer that combines feedback-driven strategy selection with adaptive mutation of paper content. By iteratively probing reviewers with dynamic perturbations, SCOPE-Fuzzer consistently uncovers vulnerabilities overlooked by static evaluation and other baselines.
Figures & tables
Figure 1: For reliability against content perturbations, aggregate results from static evaluation suggest that LLM-based reviewers can penalize perturbations, while our analysis reveals their stratified vulnerability: papers with lower initial scores are more susceptible to perturbations. Also, static evaluation is inefficient to discover effective instantiations in the perturbation space.
Table 2
Figure 2: Multi-probe evaluation with TREAP perturbations.
Figure 4
Figure 4: Fuzzing effectiveness of SCOPE and baselines.
Figure 5: Ablation of the selector and the mutator.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Framework of TriScope Review Evaluation Against Perturbation (Treap)
Figure 7: Case study for lexical syntactic complexification, verbosity increase and overclaiming.
Figure 8: Case study for logic adjustment, value alignment and empathy elicitation.
Figure 9: Human validation results across the six TREAP strategies. Top: the proportion of retained responses agreeing that the paired texts preserve the same core research content; error bars show 95% Wilson confidence intervals. Bottom: mean academic-plausibility ratings for the original and perturbed texts on a five-point scale.
Figure 10: Prompt template used for LLM-based paper review.
Review System
Mean Benign
Median
Mode
QWEN3
7.65
8
8
GPT4o-mini
7.84
8
8
SEA
5.84
6
6
Openreviewer
4.47
5
3
DeepReview
5.56
6
6
MARG
4.96
5
5
Appendix
Table 4: Statistic results of benign review scores across review systems.
Reviewer
Metric
Overclaiming
Logic Adjustment
High
Low
High
Low
QWEN3
PSRnet
-14.20
+26.09
-15.88
+29.41
PG
-0.15
0.26
-0.16
0.31
GPT-4o-mini
PSRnet
-4.72
+34.78
-3.94
+30.43
PG
-0.04
0.34
-0.03
0.30
SEA
PSRnet
-13.75
+42.37
-13.75
+42.37
Appendix
Table 5: SV under reasoning-level perturbations.
Reviewer
Metric
Value Alignment
Empathy Elicitation
High
Low
High
Low
QWEN3
PSRnet
-19.07
+25.72
-18.45
+33.33
PG
-0.20
0.27
-0.20
0.35
GPT-4o-mini
PSRnet
-3.94
+26.09
-3.15
+26.09
PG
-0.03
0.26
-0.03
0.26
SEA
PSRnet
-20.42
+36.50
-13.79
+34.92
Appendix
Table 6: SV under value-level perturbations.
PSRnet (%) ↑
Reviewer
Method
Budget
1
2
3
4
5
6
7
8
9
10
SEA
PAA
-2.15
+20.43
+25.80
+32.26
+36.56
+36.56
+37.64
+38.71
+40.86
+41.94
MPS
-5.38
+13.98
+29.03
+31.18
+36.55
+38.71
–
–
–
–
SPD
-11.83
+11.83
+22.58
+29.03
+36.56
+38.71
+39.79
+41.94
+43.01
+43.01
SCOPE
+15.05 *
+27.96 ***
+38.71
+43.01
+48.39
+51.61
+52.69
+53.76
+54.84
+58.06
Appendix
Table 7: Main fuzzing results under different budgets. PSRnet=PSR−PFR is reported in percentage points. Higher PSRnet and PG indicate stronger perturbation effects.
Reviewer
Baseline
Δ PG
padj
SEA
PAA
+0.172
<0.01
SEA
MPS
+0.215
<0.01
SEA
SPD
+0.225
<0.01
OpenReviewer
PAA
+0.150
0.213
OpenReviewer
MPS
+0.333
<0.01
OpenReviewer
SPD
+0.129
0.255
Appendix
Table 8: Paired comparison of SCOPE against the baselines on perturbation gain (PG). ΔPG=PG\textscscope−PGbaseline , where a positive value favors SCOPE. The reported effect is the mean paper-level paired PG difference. p denotes the two-sided paired permutation-test p -value across the baseline comparisons. Bold indicates a statistically significant improvement.
Figure 11: Ablation study on the heuristic selection methods.
L3S Research Center, Leibniz University Hannover Hannover, Germany · Peter L. Reichertz Institute for Medical Informatics of TU Braunschweig and Hannover Medical School Lower Saxony Center for AI and Causal Methods in Medicine (CAIMed) Hannover, Germany