cs.AIJul 1, 2026

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

Authors: Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, +2 more

Organizations: Tsinghua University, Beijing, China · Zhongguancun Academy, Beijing, China · Zhongguancun Institute of AI, Beijing, China · Huazhong University of Science and Technology, Wuhan, China · Shanghai Institute of Microsystem and Information Technology, Shanghai, China

Abstract

Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

    May 25, 2026Lingyao Li, Junjie Xiong, Changjia Zhu +5Peer ReviewDivergence

  2. Review Arcade: On the Human Alignment and Gameability of LLM Reviews

    May 27, 2026Hans Ole Hatzel, Sebastian Steindl, Jan StrichPeer ReviewReviewer

  3. Benchmarking Agentic Review Systems

    Jun 18, 2026Dang Nguyen, Wanqing Hao, Yanai Elazar +1Peer ReviewPrecision Recall