Organizations: Tsinghua University, Beijing, China · Zhongguancun Academy, Beijing, China · Zhongguancun Institute of AI, Beijing, China · Huazhong University of Science and Technology, Wuhan, China · Shanghai Institute of Microsystem and Information Technology, Shanghai, China
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.
Figures & tables
Figure 1 : LitReviewBench construction overview. Left: topic extraction from high quality AI survey papers curated from OpenAlex and assignment into an AAAI based field and subfield taxonomy. Middle: expert preference collection on LitReview Arena using blind pairwise comparisons and five dimension wise four way outcomes (D1 to D5). Right: benchmark construction by freezing arena logs into a versioned dataset for offline evaluation and standardized leaderboards.
Category
Methods
Token Cost/Query
Literature Coverage
Claim Support
Paper Structure
Research Suggestions
Overall Utility
Human
human
N/A
1787.4
1565.6
1502.5
1521.5
1668.8
GPT-5.2
38.096K
1632.8
1536.5
1322.4
1272.7
1449.1
Sonar Deep Research
322.080K
1175.9
1106.3
1262.1
1322.7
1285.9
Qwen Deep Research
70.265K
863.3
925.4
1125.0
1178.4
1117.3
OpenAI Deep Research
58.745K
882.2
943.3
857.5
867.3
874.1
Agentic Models
Average
122.297K
1138.6
1127.9
1141.8
1160.3
1181.6
Table 1: Expert-preference leaderboards on LitReviewBench. Token Cost/Query (third column): lower is better, reported in thousands of tokens based on exact measurements. Utility metrics (columns 4 to 8): higher is better, ordered as Literature Coverage, Claim Support, Paper Structure, Research Suggestions, and Overall Utility. Systems are grouped as Human, Agentic Models, and Language Models. Human achieves the highest utility scores across all metrics; GPT-5.2 leads non-human systems in Overall Utility.
Category
Methods
Literature Coverage
Claim Support
Paper Structure
Research Suggestions
Overall Utility
Human
human
378
321
317
407
310
GPT-5.2
2439
2470
2485
2367
2490
Sonar Deep Research
2012
2071
2094
1864
2096
Qwen Deep Research
1320
1320
1371
1392
1370
OpenAI Deep Research
802
767
763
775
761
Agentic Models
Average
1643.3
1657.0
1678.3
1599.5
1679.3
Table 2: Judge-induced leaderboards on LitReviewBench using Qwen/Qwen3-235B-A22B-Instruct-2507 as the evaluator. Scores are Elo ratings derived from Bradley–Terry aggregation, and higher is better.
Dimension
Judge–Expert Accuracy
Spearman’s ρ
Expert–Expert Agreement Accuracy
JudgeLLM–JudgeLLM Agreement Accuracy
D1 (Literature Coverage)
0.586
0.552
0.833
0.783
D2 (Claim Support)
0.554
0.442
0.556
0.747
D3 (Paper Structure)
0.598
0.467
0.639
0.747
D4 (Research Suggestions)
0.620
0.430
0.556
0.739
D5 (Overall Utility)
0.606
0.467
0.861
0.747
Table 3: Agreement with expert preference and reliability. Judge–expert accuracy merges Tie and Both Bad into a neutral outcome and assigns 0.5 credit regardless of the judge decision. Spearman’s ρ is the rank correlation between judge-induced and expert-induced BT and Elo leaderboards. Expert–expert agreement accuracy and JudgeLLM–JudgeLLM agreement accuracy are reported as mean pairwise accuracy. JudgeLLM–JudgeLLM agreement measures cross-model consistency between two judge models, Qwen/Qwen3-235B-A22B-Instruct-2507 and DeepSeek-V3.2, indicating whether failure patterns remain stable across different LLM judges.
Figure 2 : Expert-aligned evaluator workflow and results. (a) Context construction for the calibrated evaluator. (b) Calibration improves alignment between judge and expert votes. (c) In-domain generalization to held-out AI subfields (20%) shows robust transfer.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
D1
D2
D3
D4
D5
Human
1979.8
1955.8
1880.4
1871.7
1912.2
Open Deep Research ( LangChain AI, 2025 )
1845.6
1812.2
1868.5
1811.0
1832.5
SurveyForge ( Yan et al., 2025 )
1799.1
1775.5
1842.7
1797.2
1805.7
Agentic Models Avg.
1485.5
1468.6
1478.2
1537.2
1511.5
LLMs Avg.
1250.8
1280.5
1259.8
1255.6
1248.4
Appendix
Table 4: Specialized literature review systems outperform average model baselines but remain below human drafts.
Base Model / Method
Setting
D1
D2
D3
D4
D5
Qwen3-235B
Naive Judge
0.552
0.442
0.467
0.430
0.467
LitJudge
0.576
0.673
0.649
0.842
0.792
GPT-5.4
Naive Judge
0.673
0.370
0.721
0.766
0.770
LitJudge
0.855
0.867
0.830
0.879
0.952
Claude-Sonnet-4.5
Naive Judge
0.779
0.609
0.704
0.758
0.809
LitJudge
0.818
0.855
0.842
0.900
0.976
Appendix
Table 5: Spearman’s ρ between judge-induced and expert-induced leaderboards. LitJudge improves over naive judging across multiple base models and outperforms a random few-shot ICL control using the same number of examples, indicating that gains come from task-specific retrieval and diversity-aware calibration rather than few-shot prompting alone.
Category
D1
D2
D3
D4
D5
Human
1947.70
1742.00
2000.61
2007.63
2267.27
Agentic Models Avg.
1701.97
1612.13
1465.10
1522.75
1593.52
Language Models Avg.
1289.28
1384.32
1420.82
1384.82
1290.43
Appendix
Table 6: Biology pilot leaderboard. Agentic models include OpenAI Deep Research, Qwen Deep Research, and GPT-5.2. Language models include Claude Opus 4.5, Qwen3 235B, Grok 4, GLM 4.6, and Gemini 2.5 Pro.
Judge
D1
D2
D3
D4
D5
Naive Judge
0.7367
0.5467
0.4700
0.5133
0.5900
LitJudge
0.8833
0.6167
0.6000
0.7333
0.8833
Appendix
Table 7: LitJudge improves over the naive judge on the biology pilot study.
Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adversarial attacks remain poorly understood. We present a systematic benchmark of LLM-as-a-Reviewer on 898 papers stratified from NeurIPS and ICLR, evaluating 12 LLMs along three axes: rating calibration, divergence from human reviewers, and resistance to prompt injection embedded via an invisible font-mapping attack. We find that LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary. Prompt injection remains highly effective. Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families. While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks.
Lingyao Li, Junjie Xiong, Changjia Zhu +5
University of South Florida · Missouri University of Science and Technology · University of Alabama +3
LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to assume that not only reviewers are using LLM-assistance, but also that authors use LLMs to revise their papers before submitting. In this work, we perform empirical experiments on papers from the 2025 ACL Rolling Review (ARR) to evaluate LLM reviews from both the author and the reviewer perspective. First, we identify a limited alignment of LLM reviews with human ones. In the best-case scenario, the alignment is reasonable. However, we also find that LLM-human alignment varies substantially across prompts and models. Finally, we investigate the scenario in which the author uses an iterative draft-revise workflow to improve the submission according to the LLM review. We find that this "gaming" of LLM reviews can be effective in specific scenarios, leading to a statistically significant increase of overall scores for up to 35% of papers. We publish our code: https://github.com/uhh-hcds/reviewarcade.
Hans Ole Hatzel, Sebastian Steindl, Jan Strich
Language Technology Group, University of Hamburg, Germany · OTH Amberg-Weiden, Germany · Hub of Computing and Data Science (HCDS), University of Hamburg, Germany
A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coarse), one proprietary system (Reviewer3), and a zero-shot baseline, across six LLMs spanning frontier and efficient models. First, we study whether AI reviews on ICLR/NeurIPS papers track with papers' quality as approximated by external signals such as citations and acceptance decisions. Every system performs above chance in pairwise accuracy, and the best is OpenAIReview + GPT-5.5 at 83.0%. Second, to test whether systems can catch errors with known ground truth, we construct a perturbation benchmark that injects four categories of errors into papers across eight arXiv subject classes and measure detection recall. The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors, leaving substantial room for improvement. The union of detections across six models reaches 83.3% recall, suggesting different models detect different errors and better harness design can potentially increase performance. Beyond these benchmarks, we study a public deployment of OpenAIReview with real users. Votes on its comments skew positive at 1.44 to 1, and the most common complaints are about false positives and minor nitpicks. Together, by evaluating full review systems backed by state-of-the-art models on real research papers, we show that while AI reviews still have room for improvement, they can already track human quality judgments well, catch important errors, and earn positive feedback from real users.