Organizations: Tsinghua University, Beijing, China · Zhongguancun Academy, Beijing, China · Zhongguancun Institute of AI, Beijing, China · Huazhong University of Science and Technology, Wuhan, China · Shanghai Institute of Microsystem and Information Technology, Shanghai, China
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.
Figures & tables
Figure 1 : LitReviewBench construction overview. Left: topic extraction from high quality AI survey papers curated from OpenAlex and assignment into an AAAI based field and subfield taxonomy. Middle: expert preference collection on LitReview Arena using blind pairwise comparisons and five dimension wise four way outcomes (D1 to D5). Right: benchmark construction by freezing arena logs into a versioned dataset for offline evaluation and standardized leaderboards.
Category
Methods
Token Cost/Query
Literature Coverage
Claim Support
Paper Structure
Research Suggestions
Overall Utility
Human
human
N/A
1787.4
1565.6
1502.5
1521.5
1668.8
GPT-5.2
38.096K
1632.8
1536.5
1322.4
1272.7
1449.1
Sonar Deep Research
322.080K
1175.9
1106.3
1262.1
1322.7
1285.9
Qwen Deep Research
70.265K
863.3
925.4
1125.0
1178.4
1117.3
OpenAI Deep Research
58.745K
882.2
943.3
857.5
867.3
874.1
Agentic Models
Average
122.297K
1138.6
1127.9
1141.8
1160.3
1181.6
Table 1: Expert-preference leaderboards on LitReviewBench. Token Cost/Query (third column): lower is better, reported in thousands of tokens based on exact measurements. Utility metrics (columns 4 to 8): higher is better, ordered as Literature Coverage, Claim Support, Paper Structure, Research Suggestions, and Overall Utility. Systems are grouped as Human, Agentic Models, and Language Models. Human achieves the highest utility scores across all metrics; GPT-5.2 leads non-human systems in Overall Utility.
Category
Methods
Literature Coverage
Claim Support
Paper Structure
Research Suggestions
Overall Utility
Human
human
378
321
317
407
310
GPT-5.2
2439
2470
2485
2367
2490
Sonar Deep Research
2012
2071
2094
1864
2096
Qwen Deep Research
1320
1320
1371
1392
1370
OpenAI Deep Research
802
767
763
775
761
Agentic Models
Average
1643.3
1657.0
1678.3
1599.5
1679.3
Table 2: Judge-induced leaderboards on LitReviewBench using Qwen/Qwen3-235B-A22B-Instruct-2507 as the evaluator. Scores are Elo ratings derived from Bradley–Terry aggregation, and higher is better.
Dimension
Judge–Expert Accuracy
Spearman’s ρ
Expert–Expert Agreement Accuracy
JudgeLLM–JudgeLLM Agreement Accuracy
D1 (Literature Coverage)
0.586
0.552
0.833
0.783
D2 (Claim Support)
0.554
0.442
0.556
0.747
D3 (Paper Structure)
0.598
0.467
0.639
0.747
D4 (Research Suggestions)
0.620
0.430
0.556
0.739
D5 (Overall Utility)
0.606
0.467
0.861
0.747
Table 3: Agreement with expert preference and reliability. Judge–expert accuracy merges Tie and Both Bad into a neutral outcome and assigns 0.5 credit regardless of the judge decision. Spearman’s ρ is the rank correlation between judge-induced and expert-induced BT and Elo leaderboards. Expert–expert agreement accuracy and JudgeLLM–JudgeLLM agreement accuracy are reported as mean pairwise accuracy. JudgeLLM–JudgeLLM agreement measures cross-model consistency between two judge models, Qwen/Qwen3-235B-A22B-Instruct-2507 and DeepSeek-V3.2, indicating whether failure patterns remain stable across different LLM judges.
Figure 2 : Expert-aligned evaluator workflow and results. (a) Context construction for the calibrated evaluator. (b) Calibration improves alignment between judge and expert votes. (c) In-domain generalization to held-out AI subfields (20%) shows robust transfer.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
D1
D2
D3
D4
D5
Human
1979.8
1955.8
1880.4
1871.7
1912.2
Open Deep Research ( LangChain AI, 2025 )
1845.6
1812.2
1868.5
1811.0
1832.5
SurveyForge ( Yan et al., 2025 )
1799.1
1775.5
1842.7
1797.2
1805.7
Agentic Models Avg.
1485.5
1468.6
1478.2
1537.2
1511.5
LLMs Avg.
1250.8
1280.5
1259.8
1255.6
1248.4
Appendix
Table 4: Specialized literature review systems outperform average model baselines but remain below human drafts.
Base Model / Method
Setting
D1
D2
D3
D4
D5
Qwen3-235B
Naive Judge
0.552
0.442
0.467
0.430
0.467
LitJudge
0.576
0.673
0.649
0.842
0.792
GPT-5.4
Naive Judge
0.673
0.370
0.721
0.766
0.770
LitJudge
0.855
0.867
0.830
0.879
0.952
Claude-Sonnet-4.5
Naive Judge
0.779
0.609
0.704
0.758
0.809
LitJudge
0.818
0.855
0.842
0.900
0.976
Appendix
Table 5: Spearman’s ρ between judge-induced and expert-induced leaderboards. LitJudge improves over naive judging across multiple base models and outperforms a random few-shot ICL control using the same number of examples, indicating that gains come from task-specific retrieval and diversity-aware calibration rather than few-shot prompting alone.
Category
D1
D2
D3
D4
D5
Human
1947.70
1742.00
2000.61
2007.63
2267.27
Agentic Models Avg.
1701.97
1612.13
1465.10
1522.75
1593.52
Language Models Avg.
1289.28
1384.32
1420.82
1384.82
1290.43
Appendix
Table 6: Biology pilot leaderboard. Agentic models include OpenAI Deep Research, Qwen Deep Research, and GPT-5.2. Language models include Claude Opus 4.5, Qwen3 235B, Grok 4, GLM 4.6, and Gemini 2.5 Pro.
Judge
D1
D2
D3
D4
D5
Naive Judge
0.7367
0.5467
0.4700
0.5133
0.5900
LitJudge
0.8833
0.6167
0.6000
0.7333
0.8833
Appendix
Table 7: LitJudge improves over the naive judge on the biology pilot study.
May 27, 2026·Hans Ole Hatzel, Sebastian Steindl, Jan StrichPeer ReviewReviewer
Language Technology Group, University of Hamburg, Germany · OTH Amberg-Weiden, Germany · Hub of Computing and Data Science (HCDS), University of Hamburg, Germany