cs.CLOct 7, 2026

When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection

Authors: M. Mikail Demir, M. Abdullah Canbaz

Organizations: University at Albany, SUNY Albany, New York, USA

Abstract

Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court of Appeals decisions, we build a dataset of 3,396 parenthetical-style claims labeled as Supported, Refuted, or Not Found. We cast this task as a three-way natural language inference problem and evaluate several state-of-the-art LLMs in a zero-shot setting. Although the strongest models reach up to 0.97 accuracy, the results also show an important weakness: models still incorrectly mark unsupported claims as supported, even when the full opinion text is provided. Across models, missing support is harder to detect than direct contradiction, and fabricated but plausible citations cause the largest drop in performance. Overall, PARCEL provides a practical benchmark for testing claim-level groundedness in legal RAG systems.

Figures & tables

Explore similar work

CardsList
  1. LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

    May 11, 2026Sijia Chen, Hang Yin, Shunfan ZhouLLM EvaluationLegal IR

  2. L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

    Aug 31, 2025Boqin Yuan, Ziqi WangLegal Citation VerificationAgentic Search