Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court of Appeals decisions, we build a dataset of 3,396 parenthetical-style claims labeled as Supported, Refuted, or Not Found. We cast this task as a three-way natural language inference problem and evaluate several state-of-the-art LLMs in a zero-shot setting. Although the strongest models reach up to 0.97 accuracy, the results also show an important weakness: models still incorrectly mark unsupported claims as supported, even when the full opinion text is provided. Across models, missing support is harder to detect than direct contradiction, and fabricated but plausible citations cause the largest drop in performance. Overall, PARCEL provides a practical benchmark for testing claim-level groundedness in legal RAG systems.
Figures & tables
Figure 1 . Pipeline overview (seed extraction, synthetic claim generation and construction) of PARCEL.
Usage Type
Definition
Example
Holding rule
Summarizes the specific legal rule or final outcome established by the court. Focuses on the black letter law takeaway.
(holding that the crime of making a terroristic threat is a bail-qualifying offense)
Reasoning rationale
Explains the logical steps, policy justifications, or statutory interpretation used to reach the result. Focuses on why the court ruled that way.
(reasoning that the bail statute’s disjunctive structure creates independent categories of qualifying offenses)
Facts procedure
Highlights key factual details or procedural history, often to support an analogy or distinction. Focuses on the narrative context.
(upholding bail where the defendant made repeated calls to a police emergency line threatening to shoot officers)
Treatment relational
Describes how the decision treats another specific authority or precedent. Focuses on the relationship between cases.
(distinguishing Matter of Perlbinder holdings because that case involved conflicting provisions with different levels of specificity)
Descriptive label
A compressed, non-participial description of the proceeding or context. Used to quickly flag the nature of the case without a full sentence.
(habeas corpus proceeding converting to declaratory judgment action)
Table 1 . Usage type schema for synthetic claim generation.
Figure 2 . Distribution of adversarial generation strategies.
Model
Overall Accuracy
Macro F1
Supported
Refuted
Not Found
F1
Prec.
Rec.
F1
Prec.
Rec.
F1
Prec.
Rec.
gemini-3-flash-preview
0.97
0.97
0.95
0.91
1.00
0.98
0.99
0.96
0.97
0.99
0.95
kimi-k2-thinking
0.95
0.94
0.94
0.88
0.99
0.97
0.97
0.97
0.91
0.98
0.86
glm-4.7
0.94
0.93
0.94
0.89
0.99
0.96
0.94
0.99
0.88
0.99
0.79
qwen3-235b-a22b-instruct-2507
0.93
0.92
0.91
0.84
0.99
0.96
0.97
0.96
0.89
0.97
0.82
gemini-2.5-flash
0.90
0.90
0.87
0.77
1.00
0.93
0.96
0.91
0.89
0.99
0.81
Table 2. Model performance metrics (rounded to 2 decimals).
Model
FER (Overall)
( FER−1 ) (Refuted)
( FER0 ) (Not Found)
gemini-3-flash-preview
3.47%
3.37%
3.67%
glm-4.7
3.97%
1.18%
9.55%
kimi-k2-thinking
4.36%
2.41%
8.25%
qwen3-235b
6.17%
2.95%
12.60%
gpt-oss-120b
8.96%
6.73%
13.43%
gemini-2.5-flash
9.98%
9.14%
11.66%
Table 3 . False Entailment Rate (FER) by model.
Predicted
Ground Truth
Supported
Refuted
Not Found
Supported
97.8%
2.0%
0.2%
Refuted
4.3%
94.3%
1.4%
Not Found
15.4%
5.4%
79.2%
Table 4. Aggregate confusion matrix (all models), reported as percentage within each ground-truth class.
Model
Distr. Swap
Logic Rev.
Scope Contr.
Auth. Swap
glm-4.7
0.97
0.99
1.00
0.98
gemini-2.5-flash
0.94
0.91
0.95
0.77
gemini-3-flash-preview Preview
0.98
0.97
0.98
0.86
gpt-oss-120b
0.80
0.95
0.93
0.77
kimi-k2-thinking
0.96
0.98
0.99
0.88
llama-3.3-70b-instruct
0.81
0.98
0.97
0.93
Table 5 . Accuracy metric for refuted-claim strategies (rounded to 2 decimals).
Model
Phantom Cit.
Missing Det.
Unaddr. Arg.
glm-4.7
0.73
0.83
0.84
gemini-2.5-flash
0.80
0.97
0.66
gemini-3-flash-preview Preview
0.95
0.96
0.93
gpt-oss-120b
0.80
0.90
0.88
kimi-k2-thinking
0.86
0.91
0.82
llama-3.3-70b-instruct
0.24
0.45
0.72
Table 6 . Accuracy metric for not found-claim strategies (rounded to 2 decimals).
k (min failures)
Claims
% of 3396
k≥3
290
8.54%
k≥4
158
4.65%
k≥5
75
2.21%
k≥6
27
0.80%
k=7 (all failed)
8
0.24%
Table 7. Hard Claim Distribution. Number of claims where ≥k models failed.