When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection
Organizations: University at Albany, SUNY Albany, New York, USA
Abstract
Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court of Appeals decisions, we build a dataset of 3,396 parenthetical-style claims labeled as Supported, Refuted, or Not Found. We cast this task as a three-way natural language inference problem and evaluate several state-of-the-art LLMs in a zero-shot setting. Although the strongest models reach up to 0.97 accuracy, the results also show an important weakness: models still incorrectly mark unsupported claims as supported, even when the full opinion text is provided. Across models, missing support is harder to detect than direct contradiction, and fabricated but plausible citations cause the largest drop in performance. Overall, PARCEL provides a practical benchmark for testing claim-level groundedness in legal RAG systems.
Figures & tables
| Usage Type | Definition | Example |
|---|---|---|
| Holding rule | Summarizes the specific legal rule or final outcome established by the court. Focuses on the black letter law takeaway. | (holding that the crime of making a terroristic threat is a bail-qualifying offense) |
| Reasoning rationale | Explains the logical steps, policy justifications, or statutory interpretation used to reach the result. Focuses on why the court ruled that way. | (reasoning that the bail statute’s disjunctive structure creates independent categories of qualifying offenses) |
| Facts procedure | Highlights key factual details or procedural history, often to support an analogy or distinction. Focuses on the narrative context. | (upholding bail where the defendant made repeated calls to a police emergency line threatening to shoot officers) |
| Treatment relational | Describes how the decision treats another specific authority or precedent. Focuses on the relationship between cases. | (distinguishing Matter of Perlbinder holdings because that case involved conflicting provisions with different levels of specificity) |
| Descriptive label | A compressed, non-participial description of the proceeding or context. Used to quickly flag the nature of the case without a full sentence. | (habeas corpus proceeding converting to declaratory judgment action) |
| Model | Overall Accuracy | Macro F1 | Supported | Refuted | Not Found | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| F1 | Prec. | Rec. | F1 | Prec. | Rec. | F1 | Prec. | Rec. | |||
| gemini-3-flash-preview | 0.97 | 0.97 | 0.95 | 0.91 | 1.00 | 0.98 | 0.99 | 0.96 | 0.97 | 0.99 | 0.95 |
| kimi-k2-thinking | 0.95 | 0.94 | 0.94 | 0.88 | 0.99 | 0.97 | 0.97 | 0.97 | 0.91 | 0.98 | 0.86 |
| glm-4.7 | 0.94 | 0.93 | 0.94 | 0.89 | 0.99 | 0.96 | 0.94 | 0.99 | 0.88 | 0.99 | 0.79 |
| qwen3-235b-a22b-instruct-2507 | 0.93 | 0.92 | 0.91 | 0.84 | 0.99 | 0.96 | 0.97 | 0.96 | 0.89 | 0.97 | 0.82 |
| gemini-2.5-flash | 0.90 | 0.90 | 0.87 | 0.77 | 1.00 | 0.93 | 0.96 | 0.91 | 0.89 | 0.99 | 0.81 |
| Model | FER (Overall) | ( ) (Refuted) | ( ) (Not Found) |
|---|---|---|---|
| gemini-3-flash-preview | 3.47% | 3.37% | 3.67% |
| glm-4.7 | 3.97% | 1.18% | 9.55% |
| kimi-k2-thinking | 4.36% | 2.41% | 8.25% |
| qwen3-235b | 6.17% | 2.95% | 12.60% |
| gpt-oss-120b | 8.96% | 6.73% | 13.43% |
| gemini-2.5-flash | 9.98% | 9.14% | 11.66% |
| Predicted | |||
|---|---|---|---|
| Ground Truth | Supported | Refuted | Not Found |
| Supported | 97.8% | 2.0% | 0.2% |
| Refuted | 4.3% | 94.3% | 1.4% |
| Not Found | 15.4% | 5.4% | 79.2% |
| Model | Distr. Swap | Logic Rev. | Scope Contr. | Auth. Swap |
|---|---|---|---|---|
| glm-4.7 | 0.97 | 0.99 | 1.00 | 0.98 |
| gemini-2.5-flash | 0.94 | 0.91 | 0.95 | 0.77 |
| gemini-3-flash-preview Preview | 0.98 | 0.97 | 0.98 | 0.86 |
| gpt-oss-120b | 0.80 | 0.95 | 0.93 | 0.77 |
| kimi-k2-thinking | 0.96 | 0.98 | 0.99 | 0.88 |
| llama-3.3-70b-instruct | 0.81 | 0.98 | 0.97 | 0.93 |
| Model | Phantom Cit. | Missing Det. | Unaddr. Arg. |
|---|---|---|---|
| glm-4.7 | 0.73 | 0.83 | 0.84 |
| gemini-2.5-flash | 0.80 | 0.97 | 0.66 |
| gemini-3-flash-preview Preview | 0.95 | 0.96 | 0.93 |
| gpt-oss-120b | 0.80 | 0.90 | 0.88 |
| kimi-k2-thinking | 0.86 | 0.91 | 0.82 |
| llama-3.3-70b-instruct | 0.24 | 0.45 | 0.72 |
| (min failures) | Claims | % of 3396 |
|---|---|---|
| 290 | 8.54% | |
| 158 | 4.65% | |
| 75 | 2.21% | |
| 27 | 0.80% | |
| (all failed) | 8 | 0.24% |