cs.CLMay 20, 2026

GradeLegal: Automated Grading for German Legal Cases

Authors: Abdullah Al Zubaer, Lorenz Wendlinger, Simon Alexander Nonn, Michael Granitzer, Jelena Mitrovic

Abstract

Grading German legal exam solutions faces growing volumes and a shortage of qualified graders, delaying feedback and creating a bottleneck. At the same time, it is a high-stakes expert task, since state exam grades strongly influence career outcomes in Germany. Despite this practical relevance, literature lacks systematic studies on effective methods for grading legal exams. To address this gap, we investigate whether large language models (LLMs) can support the automated grading of German legal case solutions in criminal and public law, thereby enabling scalable feedback and student self-testing. We present a systematic evaluation of 27 proprietary and open-source LLMs, benchmarking prompting strategies that incrementally add task-related information, such as a sample solution and a grading rubric. Using quadratic weighted kappa (QWK), reasoning-oriented LLMs can approximate expert grading in public law when given a sample solution and a grading rubric (up to 0.91), compared to 0.60 in criminal law, suggesting a harder grading task in criminal law. Beyond single-model grading, ensembling improves agreement by up to 0.15 over its best member and can offer an alternative to stronger closed-source single models. In addition, our findings suggest that effective prompt design and model selection are necessary for reliable LLM-based grading of legal exams.

Explore similar work

CardsList
  1. BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

    May 27, 2026Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach +6Legal DomainGerman

  2. Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

    May 8, 2026Ramon Pires, Thales Sales Almeida, Celio Larcher Junior +6Legal DomainLlm-As-A-Judge