cs.HCJul 26, 2026

Visible to the Court: How AI Is (and Isn't) Litigated in U.S. Federal Court Opinions

Authors: Julie YuRock Yuren PangJevan HutsonKatharina Reinecke

Abstract

In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation. With courts become a recurring forum in which AI-related practices are scrutinized, it is important to empirically understand the AI litigation landscape to date. We address this gap through a systematic review of 559 U.S. federal court opinions in which AI plays a role in the parties' contentions, taxonomizing (1) common topics of dispute, (2) the AI technologies implicated, and (3) the parties involved, including common plaintiff and defendant types. We identify seven recurring dispute areas, six categories of AI technologies at the center of litigation, and four types of common litigants, alongside legal doctrines used by the litigants. A comparison of this taxonomy to the AI Incident Database revealed substantial gaps in coverage, definitions, and prevalence between documented and litigated harms, suggesting courts capture only part of the AI risk landscape. In addition, we found that court decisions primarily rely on pre-existing legal doctrines to manage AI rather than making new AI-specific laws, producing a form of "piecemeal" AI governance. As a result, federal court outcomes are shaped less by where AI has caused harms and more by which harms are cognizable under existing statutes, leading to certain AI harms remaining unresolved.

Explore similar work

May 28, 2026cs.CY

The New Pro Se: Generative AI and the Surge in Federal Civil Self-Representation

Since public access to generative AI tools became widespread, federal civil litigation has seen a marked increase in pro se (self-represented) plaintiffs. This paper analyzes that shift using ~2.8 million filings, asking whether the post-GenAI period is associated not only with more pro se filings, but also with detectable changes in complaint text, litigation outcomes, and the composition of pro se litigants. Using civil filing data from FY2008-2025, we find that the federal civil pro se plaintiff rate rose from 11.33% pre-GenAI to 16.94% post-GenAI, a 5.61 percentage-point increase that persists after trend and covariate-adjusted robustness checks. We then focus on Civil Rights and Other Statutory cases, where the increase is especially pronounced, and link case metadata to pro se complaints. Drawing on stylometric AI detection indicators, we develop an interpretable measure of AI-consistent drafting. Against a threshold calibrated to the pre-GenAI baseline, the net AI-flagged share is 13.9% of post-GenAI non-form complaints. Analysis of the AI-flagged complaints shows that they are more citation-dense, disproportionately associated with first-time rather than repeat filers, and geographically unevenly distributed. This composition pattern suggests that AI-consistent drafting is not merely a repeat-filer phenomenon; it also includes a modest, suggestive increase in name-inferred female plaintiffs. We find no evidence of improved win rates; in fact, AI-flagged complaints are more likely to be dismissed and to terminate at earlier procedural phases. These findings raise new questions about access to justice and court screening burdens, and sharpen the distinction between legal formality and legal efficacy.
Or Cohen-Sasson
Jun 19, 2026cs.CL

Who Checks the Citations? Benchmarking Legal Hallucination Detection

Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations -- with this number growing year-over-year. This study evaluates whether AI-based systems can mitigate these errors by automatically detecting hallucinations. We propose a taxonomy of legal citation hallucinations grounded in actual court filings and introduce a dataset of 1,300 brief excerpts containing injected errors. Benchmarking five models in agentic and non-agentic settings reveals that while the latest iterations perform better -- GPT-5 achieves 82.8% recall and a 60.5% F1 score in an agentic framework -- all models struggle with subtle error categories. Agentic verification remains resource-intensive, with GPT-5 averaging 16.9 steps per excerpt. Furthermore, restricted information access limits the efficacy of even the best agents. This gap creates policy concerns, as it disadvantages both AI systems and litigants who lack subscriptions to commercial legal databases. Together, our dataset, tools, and policy recommendations provide a foundation for building and auditing reliable legal citation checking tools.
Patty Liu, Dominik Stammbach, Peter Henderson
Sep 1, 2026cs.LG

The Constitutional Coverage Trilemma in AI Governance

Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of 2323 frontier LLM archetypes with a pairwise-tradeoff study of 1,6491{,}649 US participants on the same instrument, we report three facts. \emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \emph{Supply is narrow and drifting}: the 2323-archetype hull occupies 2%{\sim}2\% of the demand hull under conservative noise-matched estimation (0.10%0.10\% at full audit precision), no archetype puts helpfulness or autonomy first (37%37\% of users are constitutionally homeless), and across six model families autonomy decreases in 5/65/6, equity increases in 5/65/6, and safety increases in 4/64/6, with monotone within-family version trends (order-permutation p=0.013p = 0.013) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \emph{The fix is sparse}: a 22-vertex menu {eHON,eAUT}\{e_{\mathrm{HON}}, e_{\mathrm{AUT}}\} beats the full 2323-archetype frontier by 47%47\% on mean regret (CI [43%,52%][43\%, 52\%]); three vertex additions cut mean/worst-group regret by up to 81%81\%/64%64\%. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.
Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly +1