Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be valuable. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable. We introduce \textbf{Legal Research Bench} (LRB), a benchmark of 413 open-ended U.S. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric. We evaluate thirteen frontier models in a harness with web search, case-law search, page parsing, and retrieval tools. We score agent responses through all-pass grading with source verification, where a response is correct only if every required criterion is satisfied and its cited authorities verify. We also validate the LLM judge against expert attorneys ensuring that benchmark scores track attorney judgment. Agents remain far from reliable: among the models we tested, the strongest, Claude Opus 4.8, is fully correct on 42.9% of questions. Performance also varies substantially by task setting: all-pass rates differ across areas of law and are lower on questions requiring reconciliation of conflicting authorities. Across models, more turns, tool calls, and inference cost do not predict higher accuracy.
Figures & tables
Figure 1: Overview of the LRB evaluation setup. Each question is handed to a tool-using agent that can search the web (Tavily) and case law (CourtListener), parse and fetch pages, and recall content it has collected into a working database, before submitting a final answer. Runs are capped at three hours, and the agent manages its own errors.
Area of Law
N
Administrative / Regulatory
128
Criminal
117
Business & Commercial
88
Constitutional / Civil Rights
72
Health
58
Civil Litigation
51
Table 1: Dataset composition over all 413 questions: distribution across the eight areas of law (left) and the question-type scheme of three overlapping primary reasoning categories and two difficulty attributes (right). Questions are multi-label, so the counts sum to more than 413.
Rater
Agree
Cohen’s κ
FP
FN
Pass
GPT-5.4 (selected)
87.4
0.715
5.6
7.0
66.3
Gemini 3.1 Pro Preview
87.4
0.726
2.9
9.7
61.0
Claude Sonnet 4.6
86.5
0.708
3.2
10.3
60.7
Human (inter-rater avg)
83.4
0.644
n/a
n/a
67.7
Table 2: Judge–expert agreement vs. the human inter-rater baseline, over 341 rubric-item comparisons. Agree and κ score each judge against the attorney majority; FP and FN are the false positive and false negative rates, respectively; Pass is the share of items a rater marks pass. The Human row is pairwise attorney-to-attorney, so FP and FN are undefined, and its Pass entry is the majority’s. All three judges exceed the human agreement baseline; GPT-5.4 is selected for its balanced error profile and high agreement with the human expert majority.
Model
All-pass
Wtd pass
Sources
Turns
Tools
$/test
Claude Opus 4.8
42.9±2.4
85.4±1.1
8.6
11.7
28.7
2.97
GPT-5.5
40.0±2.4
83.4±1.1
9.0
58.5
66.8
7.34
Claude Sonnet 4.6
38.5±2.4
85.6±1.0
10.8
25.2
52.1
2.46
GLM-5.2
31.2±2.3
82.7±1.1
8.5
23.3
45.1
0.91
Gemini 3.5 Flash
30.8±2.3
83.2±1.0
7.9
47.0
47.0
1.37
MiniMax-M3
29.8±2.3
80.1±1.2
8.9
26.6
52.6
0.36
Table 3: Leaderboard over the full 413-question corpus, sorted by all-pass. Pass columns are mean ± standard error (%); Sources is the mean number of authorities the model cites ; Turns and Tools are mean agent turns and tool calls per question; $/test is mean cost per question in USD. Appendix D compares adjacent entries on paired per-question outcomes.
Figure 2: All-pass rate by area of law.
Figure 3: Pooled all-pass rate by primary reasoning category (blue) and difficulty attribute (red), with 95% confidence intervals. The reasoning categories cluster near the overall rate, whereas reconciliation is markedly harder.
Figure 4: All-pass against inference cost (left) and mean agent turns (right). The red line on the cost panel marks the Pareto frontier: the models for which no other model is both cheaper and more accurate. Tool calls are plotted separately in Appendix E , Figure 8 .
Model
All-pass (tools)
All-pass (1-shot)
Wtd (tools)
Wtd (1-shot)
Claude Opus 4.8
42.9
8.7
85.4
63.0
GPT-5.5
40.0
6.8
83.4
57.9
Claude Sonnet 4.6
38.5
7.3
85.6
58.0
GLM-5.2
31.2
3.4
82.7
50.3
Gemini 3.5 Flash
30.8
6.3
83.2
60.8
MiniMax-M3
29.8
2.9
80.1
48.4
Table 4: Ablation over the full 413-question corpus: each model run with its full tool environment versus single-shot with no tool access, sorted by all-pass with tools. All values are percentages. Removing tools lowers all-pass by 22.5 points and weighted pass by 25.6 points on average, and every tool-using model outscores every single-shot one.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Question type
Maps to
N
All-pass
95% CI
Primary reasoning categories
Statutory Interpretation
Statutory
288
25.2
[21.9, 28.5]
Application of Precedent to a Fact Pattern
Doctrinal
240
23.3
[20.1, 26.6]
Doctrinal Test Identification
Doctrinal
116
24.3
[19.5, 29.3]
Regulatory Framework Interpretation
Regulatory
91
31.0
[24.6, 37.6]
Difficulty attributes
Appendix
Table 5: The ten question types, their mapping to the category/attribute scheme used in Section 5.3 , and pooled all-pass over the full corpus (percentages, with question-level cluster-bootstrap 95% confidence intervals, which account for every question being answered by all thirteen models). N is the number of questions carrying each type.
Figure 5: Per-model all-pass rate by primary reasoning category and difficulty attribute.
Predictor
Unadjusted OR
Adjusted OR
Reconciliation
0.60
0.68
Rubric length (per criterion)
—
0.89
Source count (per source)
—
0.93
Appendix
Table 6: All-pass odds ratios for reconciliation before and after adjusting for rubric length and source count. Area-of-law indicators are included in the adjusted model as controls but are not reported. Rubric length and source count enter only the adjusted model.
Adjacent pair (all-pass %)
Discordant
McNemar p
Claude Opus 4.8 (42.9) vs. GPT-5.5 (40.0)
52/39
0.173
GPT-5.5 (40.0) vs. Claude Sonnet 4.6 (38.5)
50/44
0.536
Claude Sonnet 4.6 (38.5) vs. GLM-5.2 (31.2)
57/27
0.001
GLM-5.2 (31.2) vs. Gemini 3.5 Flash (30.8)
53/51
0.845
Gemini 3.5 Flash (30.8) vs. MiniMax-M3 (29.8)
58/54
0.705
MiniMax-M3 (29.8) vs. GLM-5.1 (27.6)
51/42
0.351
Appendix
Table 7: Adjacent-pair comparisons in leaderboard order, over the full 413-question corpus. Discordant counts are higher-ranked-only passes / lower-ranked-only passes. Bold p -values mark the two pairs that separate at the 5% level.
Figure 6: Best per-model all-pass by area of law.
Figure 7: Tool-usage composition by model. Totals at the right are mean tool calls per question.
Figure 8: All-pass against mean tool calls per question.
Figure 9: Failure decomposition by model. Each evaluated question is correct (all-pass), fails only on lower-tier [+1]/[+2] items, or is central-wrong (misses a high-value [+3] item). Central errors (right, red) make up a large share of every model’s failures; the percentage of all questions that are central-wrong is annotated.
Figure 10: Mean answer length versus all-pass per model, with the 536-word gold reference marked. Answer length is positively associated with all-pass, yet every agent at the top of the leaderboard is several times more verbose than the expert references.
Figure 11: Per-model tool-call trajectories on the sample question P-005, each call bucketed by type and shown in the order it was issued; bar length is the number of tool calls. Trajectory length varies more than tenfold and does not track all-pass performance (models are ordered top-to-bottom by corpus-wide all-pass).