cs.AIOct 5, 2026
SaveIn With the Old: Enhancing 'Classical' Document Automation with Generative AI
Organizations: Capstone Practice Systems, Harvard, Massachusetts · Maastricht Law and Tech Lab, Maastricht University, Maastricht, Netherlands
Abstract
Software-based legal assistance systems have leveraged many different forms of knowledge representation and reasoning. This article explores how document automation services rooted in expert system style and other symbolic approaches can usefully enhance and be enhanced by current generative AI approaches. We discuss the possible benefits and challenges, and report on preliminary experiments in using large language models to identify and fix issues in texts written by laypeople.
Figures & tables
Figure 1 : Screenshot of the domestic violence court order module.
| Instructions | Candidate Text | Questions (Answers) | Corrected Text |
| This is a draft contract for the sale of a good. It should adhere to the following requirements or include this information: • The name of the seller and buyer • The specific good in question • The sale price • The date of the sale • Who shall undertake the delivery • Is the good sold ”as-is” or is the seller liable for a defective good? | I am selling a car to Reed Thomas. He will pick it up on December 20th 2024. | • What is the specific car being sold? 1999 Volkswagen ’Bug.’ • What is the sale price of the car? $475 • Is the car being sold ’as-is’ or will you, as the seller, be liable for any defects? As-is • What is your name as the seller? Jane Smith | Contract for Sale of Goods This contract is entered into on December 20, 2024, by and between Jane Smith (”Seller”) and Reed Thomas (”Buyer”). 1. Description of the Good : The Seller agrees to sell, and the Buyer agrees to purchase, a 1999 Volkswagen ’Bug.’ 2. Sale Price : The agreed sale price for the vehicle is $475. 3. Date of Sale : The sale shall take place on December 20, 2024. 4. Delivery : The Buyer, Reed Thomas, shall undertake the responsibility of picking up the vehicle on the date of sale. 5. Condition of the Good : The vehicle is sold ”as-is,” and the Seller, Jane Smith, makes no warranties or guarantees regarding the condition of the vehicle. The Seller shall not be liable for any defects or issues with the vehicle after the sale is completed. By signing below, both parties agree to the terms outlined in this contract. |
Table 1 : Example of an instructions text, a candidate text, questions (LLM) with answers provided by us, and the final text written by the LLM.
| INSTRUCTIONS: |
| These are instructions on how a document should look. I will provide you with a candidate document, and you will return a list of flaws and questions to ask the user to fix these flaws. |
| Return a JSON object with the keys “questions”. The value of “questions” is a list of strings, each representing a question that should be asked to the user to cure these flaws and capture the missing information. |
| Stick to the specific instructions. Do not ask for information that is not in the instructions. Only address the flaws with one question each. If the document is a template, only ask for the variables that are in the template. |
Table 2 : Prompt for Pipeline Step 1
| INSTRUCTIONS: |
| These are instructions on how a document should look. I will provide you with a candidate document, as well as answers to follow-up questions asked to the user. Based on this information, create a new document that is correct, i.e. directly corresponds to the instructions. Return the answer as a text string. |
Table 3 : Prompt for Pipeline Step 2
| Scenario | Do the questions asked capture all of the missing elements? | Does the corrected text comply with the instructions? | Is the corrected document free of hallucinations, i.e., made up facts not part of the candidate text or questions? |
| S1 - Apartment reclamation | 100% | 100% | 100% |
| S2 - Sale Contract | 80% | 80% | 100% |
| S3 - Employment contract | 100% | 20% | 100% |
Table 4 : Evaluation results for generated questions and corrected texts. In total, 5 examples were evaluated per scenario.
| EMPLOYMENT AGREEMENT |
| This Employment Agreement, by and between {{company_name}} and {{employee_name}} is entered into as per today. As of today, {{company_name}} employs {{employee_name}}, and {{employee_name}} accepts employment, as a full-time {{job_title}}. |
Table 5 : Part of instructions for S3
Explore similar work
Legal information in India remains largely inaccessible due to the complexity of legal language and the sheer volume of legal documentation involved in research and case analysis. This paper presents NyayaAI, an AI-powered legal assistant that automates and simplifies legal workflows for lawyers, law students, and general users. The system combines Large Language Models with a Retrieval-Augmented Generation pipeline grounded in a curated Indian legal knowledge base comprising constitutional provisions, statutes, case laws, and judicial precedents. A multi-agent architecture orchestrated through the Mastra TypeScript framework coordinates a main agent with specialized sub-agents handling legal research, document summarization, case law retrieval, and drafting assistance. A compliance module validates all responses before delivery. Domain classification achieved 70% precision across test samples, with RAG retrieval precision at 74% and overall response accuracy at 72%, demonstrating that structured multi-agent LLM systems can meaningfully improve legal accessibility and workflow efficiency. The code\footnote{https://github.com/B97784/NyayaAI} is made publicly available for the benefit of the research community.
LegalCheck: Retrieval- and Context-Augmented Generation for Drafting Municipal Legal Advice Letters
Public-sector legal departments in the Netherlands face acute staff shortages, increased case volumes, and increased pressure to meet regulatory compliance. This paper presents LegalCheck, a novel system that addresses these challenges by automating the drafting of objection response letters through a combination of Retrieval-Augmented Generation (RAG) and Context-Augmented Generation (CAG). Using a large language model (LLM) alongside curated legal knowledge bases, LegalCheck performs retrieval of relevant laws and precedents, and uses controlled prompting to incorporate both external knowledge and case-specific details into a coherent draft. An expert-in-the-loop review ensures that each generated letter is legally sound and contextually appropriate. In a real-world deployment within the Municipality of Amsterdam, LegalCheck produced near-final advice letters in minutes rather than hours, while maintaining high legal consistency and factual accuracy. The output is based on actual regulations and prior cases, providing explainable outputs that captured the vast majority of required legal reasoning (often 80% to 100% of essential content). Legal professionals found that the system reduced their workload and ensured a consistent application of legal standards, without replacing human judgment. These results demonstrate substantial efficiency gains, improved legal consistency, and positive user acceptance. More broadly, this work illustrates how responsible AI can be deployed in the legal domain by augmenting LLMs with domain knowledge and governance mechanisms.
Breaking Bureaucracy: Evaluating open-source LLMs for legal document review
In this paper, we evaluate open-source generative LLMs on legal Natural Language Inference (NLI). Legal inspectorial processes take place in specific domains and often deal with confidential data. This creates a need for working with local models that do not require labeled training data. We evaluate our models on the ContractNLI benchmark and two NLI4Wills datasets. We successfully reproduce the baseline for the task (Span NLI BERT) and we evaluate multiple open-source LLMs on the same task. We analyze the invalid rate of the models, and their stability across temperature settings and domains. Among the generative models, Gemma-4 26B performs the best, reaching an accuracy of 81.2%, even outperforming the supervised model on one metric. On accuracy, it is not possible to beat the supervised model with zero-shot approaches. Qwen-3.6 35B performs well on both ContractNLI and additional datasets in the legal wills domain. Our findings indicate that zero-shot, open-source, generative LLMs are a viable alternative for real-world legal NLI when no supervised data is available. Our code is available at https://github.com/fbaratov/contractnli-llms.