CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?
Organizations: University College London · Amazon · Zhejiang University
Abstract
Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark
Figures & tables
| Trustworthy (Positive) | Untrustworthy (Negative) | ||||||
|---|---|---|---|---|---|---|---|
| Model | Accuracy | Precision | Recall | F1 | Precision | Recall | F1 |
| Sentinel | 76.60 | 73.62 | 98.69 | 84.33 | 94.23 | 37.69 | 53.85 |
| GLM-5.3 | 70.47 | 68.69 | 98.69 | 81.00 | 90.00 | 20.77 | 33.75 |
| kimi-k3 | 69.92 | 68.28 | 98.69 | 80.71 | 89.29 | 19.23 | 31.65 |
| Claude-Opus-5 | 67.13 | 66.28 | 98.69 | 79.30 | 83.33 | 11.54 | 20.27 |
| DeepSeek-V4.1-Flash | 65.74 | 65.87 | 96.07 | 78.15 | 64.00 | 12.31 | 20.65 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Location or field | Content |
|---|---|
| /testbed/task/task.json | Task JSON constructed through an explicit whitelist. |
| eval_id | Blinded task identifier. |
| repo | Repository name. |
| language | Programming language recorded in the dataset. |
| pull_request | Only base_commit , title , and body . |
| target_comment | Only body , review_path , diff_hunk , and type . |
| Parameter | Description |
|---|---|
| system_template | Receives the complete system prompt in Appendix A.1.1 as a template variable, preserving the trailing newline. |
| instance_template | Receives the fixed user prompt in Appendix A.1.2 as a template variable. |
| step_limit | Allows at most 30 model calls, including attempts that produce formatting errors and the final submission call. |
| cost_limit | A value of 0.0 disables mini-swe-agent’s monetary cost limit. The local student-inference adapter reports zero call cost; this does not imply that inference incurs no hardware cost. |
| wall_time_limit_seconds | Sets a 900-second limit per task, checked at step boundaries. |
| max_consecutive_format_errors | Stops the rollout after three consecutive formatting errors; completing a valid step resets the counter. |
| Field | Type | Content |
|---|---|---|
| teacher_private_label | Boolean | Human technical-trustworthiness label for the current training task, provided only as private teacher input. |
| student_history | Array of messages | Student-visible history before the current candidate action, including system and user instructions, earlier assistant actions, and actual tool observations. Failed calls and recovery messages are retained. |
| student_candidate_action | Assistant message or null | Candidate action proposed by the student at this state, supplied separately from the history. The initial teacher request includes this action; subsequent requests on the same correction branch set it to null . |
| Field | Type | Meaning |
|---|---|---|
| kind | "bash" , "finish" , or "needs_review" | Type of the proposed next action. |
| command | String or null | Suggested bash command. |
| prediction | Boolean or null | Final trustworthiness judgment. |
| explanation | String | Concise action rationale or explanation grounded in already observed evidence. |
| evidence | Array | Each item has the form {"message_index": integer, "quote": string} ; additional item fields are disallowed. |
| review_note | String | Private teacher note that is not exposed as the student-facing explanation. |
| kind | command | prediction | Additional requirements |
|---|---|---|---|
| bash | Nonempty command string | null | Requests the next verification step. The explanation must not assume observations that the command has not yet returned, and the command must not contain the final submission marker. |
| finish | null | Boolean | Provides an explanation and exact evidence quotations from tool messages already present in the history. |
| needs_review | null | null | Used when observed evidence conflicts with the supplied label and no reasonable additional check can resolve the conflict; the teacher does not force an answer. |
| Setting | Current implementation |
|---|---|
| Model | gpt-5.6-sol . |
| Reasoning effort | high , passed through reasoning.effort . |
| Maximum output tokens | 4096 , passed explicitly as max_output_tokens . The budget includes reasoning tokens and visible output. |
| Request timeout | 180 seconds, passed to the OpenAI SDK client; this is not a total budget for the full task. |
| SDK retries | max_retries=0 ; the SDK does not retry requests automatically. |
| API call | client.responses.create(...) . |
| Field | Content |
|---|---|
| teacher_private_label | Human label for the current training task. |
| student_history | Student-visible history before the teacher action. |
| action | Normalized teacher action, represented as an assistant message with one bash tool call. |
| execution | Actual returncode and output produced by executing the action. |
| Decision | Meaning | Effect |
|---|---|---|
| approve | The action is suitable as a teaching example. | The record is marked approved and becomes eligible for inclusion in the training set. |
| reject | The action is invalid, unsupported by its evidence, or otherwise unsuitable. | The action is not exported, and the current correction branch terminates. |
| Item | Round 1 | Round 2 |
|---|---|---|
| Action training examples | 4,622 | 4,646 |
| Training tasks represented | 701 | 714 |
| Setting | Value |
|---|---|
| Base model | Qwen/Qwen3-Coder-30B-A3B-Instruct |
| Updated parameters | LoRA parameters only; base-model weights are frozen. |
| LoRA target modules | q_proj , k_proj , v_proj , and o_proj |
| LoRA rank / alpha | 16 / 32 |
| LoRA dropout | 0.05 |
| Trainable parameters | 13,369,344 |