cs.CLSep 22, 2026

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

Authors: Felix Ringe

Abstract

Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA). Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.

Explore similar work

CardsList
  1. BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

    May 27, 2026Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach +6Legal DomainGerman

  2. Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

    May 8, 2026Ramon Pires, Thales Sales Almeida, Celio Larcher Junior +6Legal DomainLlm-As-A-Judge