Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark
Figures & tables
Figure 1: Given a pull request and a generated review comment, a judge agent inspects the relevant software context before deciding whether the comment is technically trustworthy.
Figure 2: Overview of the CRJudgeBench construction pipeline.
Figure 3
Figure 4: Teacher guidance after an inconclusive code search. Before the student concludes that the reviewed code is absent, the teacher proposes inspecting the code patch.
Trustworthy (Positive)
Untrustworthy (Negative)
Model
Accuracy
Precision
Recall
F1
Precision
Recall
F1
Sentinel
76.60
73.62
98.69
84.33
94.23
37.69
53.85
GLM-5.3
70.47
68.69
98.69
81.00
90.00
20.77
33.75
kimi-k3
69.92
68.28
98.69
80.71
89.29
19.23
31.65
Claude-Opus-5
67.13
66.28
98.69
79.30
83.33
11.54
20.27
DeepSeek-V4.1-Flash
65.74
65.87
96.07
78.15
64.00
12.31
20.65
Table 2: Overall accuracy and class-specific precision, recall, and F1 score on the 359-instance CRJudgeBench test split. All models are incorporated by mini-swe-agent.
Figure 5: Repository-level accuracy on the CRJudgeBench test split.
Table 7
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Location or field
Content
/testbed/task/task.json
Task JSON constructed through an explicit whitelist.
eval_id
Blinded task identifier.
repo
Repository name.
language
Programming language recorded in the dataset.
pull_request
Only base_commit , title , and body .
target_comment
Only body , review_path , diff_hunk , and type .
Appendix
Table 5: Files and fields visible to the student agent.
Parameter
Description
system_template
Receives the complete system prompt in Appendix A.1.1 as a template variable, preserving the trailing newline.
instance_template
Receives the fixed user prompt in Appendix A.1.2 as a template variable.
step_limit
Allows at most 30 model calls, including attempts that produce formatting errors and the final submission call.
cost_limit
A value of 0.0 disables mini-swe-agent’s monetary cost limit. The local student-inference adapter reports zero call cost; this does not imply that inference incurs no hardware cost.
wall_time_limit_seconds
Sets a 900-second limit per task, checked at step boundaries.
max_consecutive_format_errors
Stops the rollout after three consecutive formatting errors; completing a valid step resets the counter.
Appendix
Table 6: mini-swe-agent parameters used to collect the first-iteration student trajectories.
Field
Type
Content
teacher_private_label
Boolean
Human technical-trustworthiness label for the current training task, provided only as private teacher input.
student_history
Array of messages
Student-visible history ht before the current candidate action, including system and user instructions, earlier assistant actions, and actual tool observations. Failed calls and recovery messages are retained.
student_candidate_action
Assistant message or null
Candidate action proposed by the student at this state, supplied separately from the history. The initial teacher request includes this action; subsequent requests on the same correction branch set it to null .
Appendix
Table 7: Top-level fields in the teacher input payload.
Field
Type
Meaning
kind
"bash" , "finish" , or "needs_review"
Type of the proposed next action.
command
String or null
Suggested bash command.
prediction
Boolean or null
Final trustworthiness judgment.
explanation
String
Concise action rationale or explanation grounded in already observed evidence.
evidence
Array
Each item has the form {"message_index": integer, "quote": string} ; additional item fields are disallowed.
review_note
String
Private teacher note that is not exposed as the student-facing explanation.
Appendix
Table 8: Top-level fields in a teacher proposal.
kind
command
prediction
Additional requirements
bash
Nonempty command string
null
Requests the next verification step. The explanation must not assume observations that the command has not yet returned, and the command must not contain the final submission marker.
finish
null
Boolean
Provides an explanation and exact evidence quotations from tool messages already present in the history.
needs_review
null
null
Used when observed evidence conflicts with the supplied label and no reasonable additional check can resolve the conflict; the teacher does not force an answer.
Appendix
Table 9: Semantic constraints for teacher action types.
Setting
Current implementation
Model
gpt-5.6-sol .
Reasoning effort
high , passed through reasoning.effort .
Maximum output tokens
4096 , passed explicitly as max_output_tokens . The budget includes reasoning tokens and visible output.
Request timeout
180 seconds, passed to the OpenAI SDK client; this is not a total budget for the full task.
SDK retries
max_retries=0 ; the SDK does not retry requests automatically.
API call
client.responses.create(...) .
Appendix
Table 10: Teacher API and orchestration configuration.
Field
Content
teacher_private_label
Human label for the current training task.
student_history
Student-visible history before the teacher action.
action
Normalized teacher action, represented as an assistant message with one bash tool call.
execution
Actual returncode and output produced by executing the action.
Appendix
Table 11: Fields supplied to the reviewer.
Decision
Meaning
Effect
approve
The action is suitable as a teaching example.
The record is marked approved and becomes eligible for inclusion in the training set.
reject
The action is invalid, unsupported by its evidence, or otherwise unsuitable.
The action is not exported, and the current correction branch terminates.
Appendix
Table 12: Reviewer decisions and their effects on the training pipeline.
Item
Round 1
Round 2
Action training examples
4,622
4,646
Training tasks represented
701
714
Appendix
Table 13: Action-level training data used in the two rounds.
Setting
Value
Base model
Qwen/Qwen3-Coder-30B-A3B-Instruct
Updated parameters
LoRA parameters only; base-model weights are frozen.
LoRA target modules
q_proj , k_proj , v_proj , and o_proj
LoRA rank / alpha
16 / 32
LoRA dropout
0.05
Trainable parameters
13,369,344
Appendix
Table 14: Action-level distillation training configuration.
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as evaluators in four tasks: picking their own solution from a pair, judging whether a single solution is their own, identifying which of two solutions a named model wrote, and judging quality blind. In the single-solution task, balanced accuracy is 49-58% for all 15 model-benchmark combinations, while raw accuracy (38-67%) mostly reflects how readily a model claims authorship. In the pairwise task, accuracy across 14 evaluator-opponent combinations correlates at r=0.93 with how often the evaluator's solution is longer. Attribution to a named model succeeds on some pairs and is consistently inverted on others. A rule-based normalization that strips docstrings, comments, type hints, and local names preserves Pass@1 and leaves ten of twelve re-tested results at chance; the other two follow a length difference it leaves, although a trained classifier still separates most normalized pairs. Claude Haiku's self-preference also disappears. We recommend reporting balanced accuracy, heuristic baselines, and label consistency.
Large Language Models are increasingly used as judges to evaluate code artifacts when exhaustive human review or executable test coverage is unavailable. LLM-judge is increasingly relevant in agentic software engineering workflows, where it can help rank candidate solutions and guide patch selection. While attractive for scale, current practice lacks a principled account of reliability and bias: repeated evaluations of the same case can disagree; small prompt edits can swing outcomes; and seemingly semantics-preserving, human-equivalent perturbations may elicit divergent verdicts. This paper studies LLM-as-a-Judge for code through a measurement-first lens. We analyze two pointwise judging regimes across code generation, code repair task, and test generation, and we systematically probe prompt-induced biases. Our study considers difficulty levels for repeated runs and controlled prompt interventions that isolate one presentation cue at a time, and it evaluates judges using consistency and sensitivity to bias. We find that judge decisions are highly sensitive to prompt biases even when the underlying code snippet is unchanged. Across all three tasks, several biases systematically shift preferences toward the option favored by the prompt, improving accuracy when that option aligns with the gold answer but substantially reducing it otherwise. In some settings, these effects are large enough to change task-level conclusions and alter relative model rankings. These findings show that reported judge performance may reflect prompt artifacts rather than stable assessment ability, posing a direct threat to the validity and reproducibility of code evaluation. We therefore argue that LLM-as-a-Judge studies should report bias sensitivity alongside accuracy and incorporate explicit controls to support more trustworthy model comparison in software engineering.
Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is merged into shared repositories. However, it is unclear whether review agents can detect vulnerability-introducing code when an attacker controls both the code change and the persuasive Pull Request (PR) narrative designed to mask it. We introduce SEVRA-BENCH (Social Engineering of Vulnerabilities in Review Agents), a benchmark that measures how often a review agent approves such adversarial PR s. Each PR in SEVRA-BENCH is built from a historical commit that fixed a vulnerability. We automatically reverse that fix to extract the original vulnerable code, and submit the resulting code change as a PR wrapped in one of 15 social-engineering framings. To test review-agent resilience to narrative manipulation, these framings vary dimensions such as supporting evidence, conveyed urgency, signals of prior approval, and appeals to authority. SEVRA-BENCH evaluates a retained challenge split of roughly 1000 adversarial PRs drawn from publicly disclosed vulnerability fixes across the top 10 entries of the MITRE's 2025 most dangerous software weaknesses. Evaluating 8 review agents against this benchmark, we reveal that review agents are susceptible to narrative manipulation, exposing a significant gap in security capabilities.
Rui Melo, Riccardo Fogliato, Sean Zhou +2
1Carnegie Mellon University · 2Microsoft Core AI · 3Independent Researcher +1