cs.AISep 29, 2026

Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers

Authors: Zhuo Chen, Hao Zeng, Jiawei Liu, Guoxiu He, Le Cai, Liu Haotan, Li Wenbo, Yong Huang, +1 more

Organizations: Wuhan University · East China Normal University

Abstract

The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Experiments on representative LLM-based reviewers reveal two limitations of static evaluation: stratified vulnerability, where perturbation effects depend on whether the paper's original review score is high or low, and perturbation undercoverage, where a single template misses vulnerabilities exposed by diverse realizations. To address these limitations, we propose SCOPE-Fuzzer, a strategy-aware fuzzer that combines feedback-driven strategy selection with adaptive mutation of paper content. By iteratively probing reviewers with dynamic perturbations, SCOPE-Fuzzer consistently uncovers vulnerabilities overlooked by static evaluation and other baselines.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

    May 25, 2026Lingyao Li, Junjie Xiong, Changjia Zhu +5Peer ReviewDivergence

  2. LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

    Jun 23, 2026Thi Huyen Nguyen, Zahra AhmadiPeer ReviewLarge Language Model Evaluation

  3. SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

    Apr 29, 2026Yuan Xin, Yixuan Weng, Minjun Zhu +5Peer ReviewAdversarial Prompts