cs.CLOct 5, 2026

Breaking Bureaucracy: Evaluating open-source LLMs for legal document review

Authors: Farrukh Baratov, Niki van Stein, Suzan Verberne

Organizations: Leiden University Leiden, the Netherlands

Abstract

In this paper, we evaluate open-source generative LLMs on legal Natural Language Inference (NLI). Legal inspectorial processes take place in specific domains and often deal with confidential data. This creates a need for working with local models that do not require labeled training data. We evaluate our models on the ContractNLI benchmark and two NLI4Wills datasets. We successfully reproduce the baseline for the task (Span NLI BERT) and we evaluate multiple open-source LLMs on the same task. We analyze the invalid rate of the models, and their stability across temperature settings and domains. Among the generative models, Gemma-4 26B performs the best, reaching an accuracy of 81.2%, even outperforming the supervised model on one metric. On accuracy, it is not possible to beat the supervised model with zero-shot approaches. Qwen-3.6 35B performs well on both ContractNLI and additional datasets in the legal wills domain. Our findings indicate that zero-shot, open-source, generative LLMs are a viable alternative for real-world legal NLI when no supervised data is available. Our code is available at https://github.com/fbaratov/contractnli-llms.

Figures & tables

Appendix figures & tables54 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks

    Jun 8, 2026Yifan Chen, Haitao Li, Yiran Hu +6Legal Reasoning TasksTask-Specific Rubrics

  2. Know Your Limits : On the Faithfulness of LLMs as Solvers and Autoformalizers in Legal Reasoning

    Jun 15, 2026Olivia Peiyu Wang, Sanna Wong-Toropainen, Daneshvar Amrollahi +4Legal Reasoning TasksLLM Reasoning Strategies

  3. BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

    May 27, 2026Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach +6Legal Reasoning TasksGerman