cs.SESep 29, 2026

CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?

Authors: Yue Pan, Jiawei Li, Ziyuan Zhang, Xiangxin Zhao, He Ye

Organizations: University College London · Amazon · Zhejiang University

Abstract

Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

    Sep 24, 2026Ehsan Barkhordar, Surendrabikram ThapaCode QualityMle-Bench Lite

  2. Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering

    Apr 18, 2026Zixiao Zhao, Amirreza Esmaeili, Fatemeh FardLlm-As-A-JudgeSoftware Engineering

  3. SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

    Jun 11, 2026Rui Melo, Riccardo Fogliato, Sean Zhou +2Code QualityReviewer