Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
Organizations: Sakana AI · National University of Singapore
Abstract
The rapid growth of scientific publishing has strained peer review, particularly in machine learning, raising concerns about declining review quality and increasing reviewer workload. Large language models (LLMs) have been proposed as automated review assistants, yet their evaluation has focused largely on imitating human-written reviews rather than supporting the core functions of peer review. Here, we introduce a verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task. We present a scalable benchmark that evaluates review systems' ability to identify logical contradictions, constructed through synthetic insertion of errors into conference papers, yielding unambiguous evaluation targets and enabling systematic comparison. We further propose a Multi-Layered Review (MLR) framework that prioritizes detailed manuscript comprehension before review generation, aligning more closely with human reviewing practices while improving token efficiency. Across evaluations, our approach demonstrates strong alignment with human review scores, achieves high error detection performance, and provides complementary perspectives on reviewer focus. These improvements can be attributed to both the choice of the underlying LLM and the design of our system. At the same time, we corroborate persistent vulnerabilities to adversarial manipulation, underscoring the need for robustness in automated review systems. Our findings highlight the importance of rigorous, error-focused evaluation to guide responsible deployment of LLM-based tools in peer review and other critical scientific workflows.
Figures & tables
| Conference | Year | # Papers | Node Distance Range | # Contradictions |
| ACL | 2025 | 51 | 0-8 | 242 |
| AISTATS | 2025 | 51 | 0-7 | 209 |
| CVPR | 2025 | 51 | 0-7 | 240 |
| ICML | 2025 | 51 | 0-7 | 238 |
| NeurIPS | 2024 | 53 | 0-7 | 235 |
| Node Type | Subtype | Definition |
| Claim | Main | High-level assertions that support the core contributions |
| Secondary | Assertions that support main claims, but are not key results | |
| Tertiary | Low-impact statements that supplement secondary claims | |
| Evidence | High | Strong results or empirical observations supporting a claim |
| Medium | Supporting findings that reinforce a claim but are not crucial | |
| Low | Minor observations with limited influence on conclusions |
| Setting / Method | MLR (Ours) | LLM-Review | AI Reviewer | AgentReview |
| Similar | 26.07 | 5.21 | 13.74 | 18.48 |
| Exact | 16.11 | 2.37 | 9.00 | 5.69 |
| Target | Aspect |
| Overall Motivation | Communication Clarity |
| Method | Validity |
| Theory | Novelty |
| Experiment | Impact |
| Conclusion | Not-specific |
| Paper |
| Model/System | ST | SA | S-Avg | WT | WA | W-Avg | Avg |
| GPT-4o mini* | 0.083 | 0.047 | 0.065 | 0.054 | 0.441 | 0.248 | 0.156 |
| GPT-5 mini | 0.202 | 0.109 | 0.156 | 0.227 | 0.190 | 0.209 | 0.182 |
| GPT-5.1 | 0.149 | 0.058 | 0.104 | 0.108 | 0.491 | 0.300 | 0.202 |
| Gemini 2.5 Pro | 0.052 | 0.069 | 0.061 | 0.142 | 0.171 | 0.157 | 0.108 |
| Gemini 3 Pro | 0.207 | 0.121 | 0.164 | 0.116 | 0.091 | 0.104 | 0.134 |
| Claude Sonnet 4.5 | 0.142 | 0.107 | 0.125 | 0.047 | 0.090 | 0.069 | 0.097 |
| Conference | Year | Decision | # Papers |
| ICML | 2025 | Accept | 51 |
| Reject | 50 | ||
| ICLR | 2025 | Accept | 50 |
| Reject | 50 | ||
| NeurIPS | 2024 | Accept | 52 |
| Reject | 50 |
| Venue | Method | Pearson | Spearman | Kendall |
| NeurIPS 2024 | Human (Reference) | 0.781 | 0.700 | 0.547 |
| MLR (Ours) | 0.451 | 0.386 | 0.299 | |
| LLM-Review | 0.358 | 0.457 | 0.377 | |
| AI Reviewer | 0.328 | 0.331 | 0.265 | |
| AgentReview | 0.167 | 0.139 | 0.103 | |
| ICML 2025 | Human (Reference) | 0.684 | 0.599 | 0.477 |
| Method | LLM-judge | Self | ||||
| Before | After | Before | After | |||
| MLR (Ours) | \textbf{0.70}_{{\color[rgb]{0,0,0}\pm{2.46}}} | \textbf{0.26}_{{\color[rgb]{0,0,0}\pm{1.85}}} | ||||
| LLM-Review | 2.42_{{\color[rgb]{0,0,0}\pm{1.47}}} | - | - | - | ||
| AI Reviewer | \underline{0.72}_{{\color[rgb]{0,0,0}\pm{0.98}}} | 1.32_{{\color[rgb]{0,0,0}\pm{1.09}}} | ||||
| AgentReview | 1.68_{{\color[rgb]{0,0,0}\pm{1.09}}} | \underline{0.43}_{{\color[rgb]{0,0,0}\pm{0.69}}} | ||||
| Method | Token count | Unit price | Cost ($) | |||
| ($/1M tokens) | ||||||
| Input | Output | Input | Output | Input | Output | |
| MLR (Ours) | 189,062 | 3,913 | 3 | 15 | 0.42 | 0.05 |
| Appendix | 66,521 | 1,011 | 0.8 | 4 | 0.05 | |
| Review | 122,541 | 2,902 | 3 | 15 | 0.37 | 0.04 |
| Literature Review (opt.) | 312,767 | 2,032 | 3 | 15 | 0.94 | 0.03 |
| Section | Reported | Agreed Comments |
| Comments | (Comment Acceptance Rate) | |
| Overall Recommendation | 31 | 29 (94%) |
| Strengths | 90 | 83 (92%) |
| Questions | 68 | 54 (79%) |
| To-Do List | 107 | 83 (78%) |
| Weaknesses | 82 | 56 (68%) |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Outline | Review |
| Gemini 2.5 Pro | 72.40 | 93.11 |
| Gemini 3 Pro | 99.87 | 99.92 |
| GPT-5.2 | 88.77 | 96.35 |
| GPT-4.1 | 92.63 | 100.00 |
| o3 | 97.42 | 99.90 |
| o4-mini | 94.70 | 99.52 |
| Evaluation/System | Appendix | Literature Review | Review |
| Contradiction Benchmark | ✓ | ||
| WithdrarXiv-Check | ✓ | ||
| Explicit manipulation | ✓ | ||
| Conference submissions | ✓ | ✓(Web search) | ✓ |
| Focus distributions | ✓ | ✓(Web search) | ✓ |
| User study | ✓ | ✓ | ✓ |
| Target | Definition |
| Overall Motivation | Significance of challenges the paper wants to address |
| Method | Approach, artifact, or solution the paper uses to address the problem |
| Theory | Theoretical components, claims, and logic of the paper |
| Experiment | Evaluation of the effectiveness and validity of the method |
| Conclusion | Discussion, insights, and takeaways |
| Paper | General comments or multiple aspects |
| Method | Node Distance | |||||||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Full | |
| MLR (1 review) | 60.79 | 32.43 | 19.56 | 14.47 | 12.85 | 4.17 | 16.67 | 0.00 | 0.00 | 28.32 |
| MLR (4 reviews) | 73.43 | 47.17 | 31.67 | 27.41 | 21.46 | 15.21 | 38.89 | 0.00 | 0.00 | 40.95 |
| LLM-Review (Claude) | 35.40 | 22.75 | 7.77 | 7.68 | 6.99 | 8.33 | 0.00 | 0.00 | 0.00 | 16.43 |
| LLM-Review | 14.56 | 7.89 | 3.63 | 3.03 | 2.20 | 2.29 | 0.00 | 0.00 | 0.00 | 6.39 |
| AI Reviewer | 11.17 | 7.73 | 6.89 | 3.73 | 1.71 | 2.08 | 0.00 | 14.00 | 0.00 | 6.50 |
| Page count/ | 1–10 | 11–20 | 21–30 | Total | ||||||||
| Subject | # Paper | Similar | Exact | # Paper | Similar | Exact | # Paper | Similar | Exact | # Paper | Similar | Exact |
| math | 37 | 35.1 | 21.6 | 46 | 21.7 | 17.4 | 22 | 13.6 | 9.1 | 105 | 24.8 | 17.1 |
| cs | 9 | 66.7 | 22.2 | 15 | 26.7 | 13.3 | 7 | 14.3 | 14.3 | 31 | 35.5 | 16.1 |
| physics | 10 | 40.0 | 20.0 | 5 | 80.0 | 80.0 | 1 | 0.0 | 0.0 | 16 | 50.0 | 37.5 |
| cond-mat 6 6 6 Condensed Matter | 10 | 0.0 | 0.0 | 2 | 50.0 | 0.0 | 3 | 33.3 | 0.0 | 15 | 13.3 | 0.0 |
| Others | 24 | 25.0 | 16.7 | 11 | 9.1 | 0.0 | 9 | 11.1 | 11.1 | 44 | 18.2 | 11.4 |
| Venue | Method | Pearson | Spearman | Kendall |
| NeurIPS 2024 | Human (Reference) | 0.781 | 0.700 | 0.547 |
| MLR-Combined | 0.395 | 0.355 | 0.282 | |
| MLR | 0.451 | 0.386 | 0.299 | |
| LLM-Review | 0.358 | 0.457 | 0.377 | |
| AI Reviewer | 0.328 | 0.331 | 0.265 | |
| AgentReview | 0.167 | 0.139 | 0.103 |
| Decision | Venue | Method | Pearson | Spearman | Kendall |
| Accepted | NeurIPS 2024 | Human (Reference) | 0.677 | 0.657 | 0.497 |
| MLR-Combined | 0.042 | 0.032 | 0.026 | ||
| MLR | 0.060 | 0.038 | 0.020 | ||
| LLM-Review | 0.195 | 0.220 | 0.185 | ||
| AI Reviewer | -0.004 | -0.012 | -0.013 | ||
| AgentReview | 0.049 | 0.019 | 0.008 |
| Venue | Method | Pearson | Spearman | Kendall |
| NeurIPS 2024 | MLR-Combined | 0.406 | 0.363 | 0.297 |
| MLR | 0.432 | 0.384 | 0.311 | |
| AI Reviewer | 0.445 | 0.444 | 0.350 | |
| AgentReview | 0.066 | 0.083 | 0.064 | |
| ICML 2025 | MLR-Combined | 0.517 | 0.386 | 0.333 |
| MLR | 0.445 | 0.305 | 0.258 |