JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction
Authors: Zhaolu Kang, Yantao Liu, Tailong Luo, Leqi Zheng, Lei Wei, Chenghua Zhu, Junhao Gong, Jiachen Qian, +9 more
Organizations: Hunyuan, Tencent · Peking University · Tsinghua University · City University of Hong Kong · University of California · University of Illinois Urbana-Champaign · Zhejiang University · University of Hong Kong
Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, and while some have attempted to evaluate reasoning quality, their evaluations are indirect, often relying on LLM-generated rubrics that reflect model-internal preferences rather than the inherent logical structure of legal adjudication. We propose Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction. JPO first uses teacher-generated rationales to supervise a standardized four-step reasoning process, and then applies reinforcement learning with a composite reward over legal prediction quality, reasoning structure completeness, and cross-step consistency. JPO further introduces token-level advantage reweighting and adaptive clipping for legally salient reasoning segments. Experiments on multiple open-source language models and three Chinese legal benchmarks show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.
Figures & tables
Figure 1: Overview of JPO. Structured SFT with teacher-generated four-step rationales, followed by reinforcement learning with rewards for legal accuracy, reasoning completeness, and cross-step consistency.
Method
Set.
JPO-Dataset
CAIL2018
LawBench
Art.
Charge
Sent.
4-Step
Full
Art.
Charge
Sent.
4-Step
Full
Art.
Charge
Sent.
4-Step
Full
(F1)
(F1)
(Score)
(Comp.)
(Chain)
(F1)
(F1)
(Score)
(Comp.)
(Chain)
(F1)
(F1)
(Score)
(Comp.)
(Chain)
Open-Source Results
Qwen3-4B-Instruct
Pre-trained
0.521
0.468
0.174
0.315
0.208
0.485
0.442
0.158
0.291
0.187
0.463
0.419
0.142
0.264
0.175
SFT
0.884
0.858
0.405
0.902
0.652
0.856
0.831
0.381
0.882
0.627
0.834
0.809
0.364
0.857
0.598
JPO
0.931
0.916
0.542
0.966
0.789
0.903
0.888
0.514
0.951
0.758
0.872
0.861
0.485
0.935
0.724
Table 1: Main results on JPO-Dataset, CAIL2018, and LawBench. We report article F1, charge F1, sentence score, 4-Step Completeness, and Full-Chain Consistency.
Dataset
SFT Train
RL Train
Test
JPO-Dataset
239,515
9,691
20,396
CAIL2018
–
–
30,000
LawBench
–
–
1,500
Table 2: Dataset statistics used in our experiments.
Figure 2: Reasoning-chain breakdown on two representative 3B backbones. F → A, A → C, and C → S denote fact-to-article, article-to-charge, and charge-to-sentence consistency, respectively.
Figure 3: Representative qualitative example illustrating the two-stage process of JPO on a criminal case.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Value
Time range of collected cases
2024–2026
Number of unique charges
192
Number of unique statutory articles
176
Average fact length (tokens)
217.1
Median fact length (tokens)
194
Average number of articles per case
1.04
Appendix
Table 3: Profile statistics of JPO-Dataset.
Figure 4: Training dynamics of JPO on the representative Qwen2.5-3B backbone. The SFT stage mainly improves output validity and coarse prediction quality, while the RL stage yields the largest gains in sentence prediction and reasoning consistency.
Method
Art. F1
Charge F1
Sent.
Full-Chain
JPO
0.929
0.921
0.536
0.791
w/o SFA
0.923
0.915
0.520
0.768
w/o SAC
0.920
0.908
0.513
0.755
w/o SCS
0.925
0.916
0.497
0.763
Entropy-only weighting
0.923
0.912
0.523
0.769
Logic-only weighting
0.924
0.914
0.525
0.771
Appendix
Table 4: Extended ablation analysis on Qwen2.5-3B. Bold indicates the full JPO model, and underlined text indicates key ablations.
α
β
γ
Art. F1
Charge F1
Sent.
0.85
0.05
0.10
0.923
0.913
0.518
0.75
0.0625
0.1875
0.929
0.921
0.536
0.65
0.0625
0.2875
0.926
0.917
0.530
0.33
0.33
0.33
0.904
0.893
0.494
Appendix
Table 5: Sensitivity to reward weight composition.
ψ
Art. F1
Charge F1
Sent.
0.0
0.922
0.911
0.519
0.3
0.924
0.915
0.523
0.6
0.929
0.921
0.536
0.9
0.927
0.918
0.531
Appendix
Table 6: Sensitivity to token-level advantage scaling ψ .
ζ
Art. F1
Charge F1
Sent.
0.0
0.924
0.914
0.525
0.25
0.926
0.917
0.529
0.5
0.929
0.921
0.536
0.75
0.925
0.916
0.527
1.0
0.923
0.912
0.523
Appendix
Table 7: Sensitivity to entropy–logic mixing coefficient ζ .
Figure 5: Hyper-parameter sensitivity of JPO on the representative Qwen2.5-3B backbone.
Method
Set.
JPO-Dataset
CAIL2018
LawBench
Art.
Charge
Sent.
4-Step
Full
Art.
Charge
Sent.
4-Step
Full
Art.
Charge
Sent.
4-Step
Full
(F1)
(F1)
(Score)
(Comp.)
(Chain)
(F1)
(F1)
(Score)
(Comp.)
(Chain)
(F1)
(F1)
(Score)
(Comp.)
(Chain)
DeepSeek-V3.2
Zero-Shot
0.924
0.908
0.512
0.958
0.771
0.896
0.875
0.481
0.942
0.738
0.864
0.841
0.439
0.925
0.706
Qwen3-32B
Zero-Shot
0.917
0.911
0.504
0.961
0.759
0.887
0.882
0.473
0.945
0.722
0.858
0.835
0.428
0.931
0.694
GPT-5.2
Zero-Shot
0.912
0.894
0.489
0.947
0.744
0.879
0.861
0.455
0.933
0.705
0.845
0.829
0.412
0.916
0.683
Claude-Sonnet-4.5
Zero-Shot
0.908
0.891
0.493
0.952
0.751
0.882
0.866
0.462
0.937
0.713
0.849
0.824
0.417
0.918
0.687
Appendix
Table 8: Zero-shot performance of proprietary models. These results are for reference only and are not directly comparable to our open-source models trained with task-specific post-training.
Method
Paradigm
Art. F1
Charge F1
Sentencing (as reported)
TopJudge
dependency-aware
0.737
0.800
0.357 (Acc)
MPBFN
dependency-aware
0.706
0.757
0.362 (Acc)
LADAN
structure & knowledge
0.738
0.801
0.361 (Acc)
NeurJudge
structure & knowledge
0.797
0.807
0.374 (Acc)
CTM
structure & knowledge
0.768
0.780
0.374 (Acc)
EPM
dependency-aware
0.781
0.814
0.367 (Acc)
Appendix
Table 9: Comparison with SOTA LJP methods on CAIL2018.
System
Fact
Article
Charge
Sentence
SFT (Qwen2.5-7B-Instruct)
0.73
0.63
0.58
0.49
JPO (Qwen2.5-7B-Instruct)
0.87
0.83
0.79
0.71
Improvement
+0.14
+0.20
+0.21
+0.22
Appendix
Table 10: Pilot blind expert evaluation results on 400 randomly sampled cases from the JPO-Dataset test set.
Teacher
Stage
Art. (F1)
Charge (F1)
Sent. (Score)
Qwen2.5-72B-Instruct
SFT
0.893
0.871
0.422
JPO
0.937
0.928
0.551
DeepSeek-V2
SFT
0.901
0.875
0.438
JPO
0.939
0.926
0.554
Appendix
Table 11: Teacher model sensitivity analysis on JPO-Dataset.
Model
ξ
Pre-trained
SFT
JPO
Qwen2.5-3B-Instruct
1.0
0.317
0.613
0.747
2.0
0.185
0.472
0.614
3.0
0.114
0.391
0.536
4.0
0.086
0.347
0.475
5.0
0.067
0.305
0.437
Qwen2.5-7B-Instruct
1.0
0.430
0.657
0.765
Appendix
Table 12: Sentence score sensitivity to ξ on JPO-Dataset.
Method
GPU Hours
Avg. Length
Std. Full-Chain
SFT
34
456
0.018
Vanilla PPO
49
471
0.025
JPO
53
478
0.015
Appendix
Table 13: Training efficiency and stability on Qwen2.5-3B.
Error Type
Proportion
Missing secondary article
27.5%
Correct charge but wrong sentence band
25.3%
Confusion between neighboring charges
21.6%
Incomplete mitigation/aggravation reasoning
16.7%
Formatting or parsing errors
8.9%
Appendix
Table 14: Main remaining error types of JPO.
Automatic proxy
Expert step judged
Spearman ρ
SFA (fact-to-article)
Statutory analysis
0.67
SAC (article-to-charge)
Charge determination
0.71
SCS (charge-to-sentence)
Sentence prediction
0.64
Full-Chain Consistency
Overall reasoning
0.72
Appendix
Table 15: Correlation between the automatic consistency proxies and blind expert step-level ratings, computed on the 800-case expert study of Appendix O.6 (Qwen2.5-7B-Instruct).
Reasoning structure
Art. F1
Charge F1
Sent.
Full-Chain
Standard CoT (free-form)
0.921
0.907
0.503
0.724
Fact–Element–Charge (3-step)
0.928
0.914
0.517
0.749
Issue-tree decomposition
0.931
0.918
0.525
0.768
Four-step F → A → C → S
0.937
0.928
0.551
0.806
Appendix
Table 16: Reasoning-structure comparison under the JPO pipeline (Qwen2.5-7B-Instruct, JPO-Dataset). Rewards are adapted to each structure; all other settings are identical.
Evaluation subset
Model
Art. F1
Charge F1
Sent.
Full-Chain
Single-def./single-charge
SFT
0.893
0.871
0.422
0.681
JPO
0.937
0.928
0.551
0.806
Multi-charge ( n =1,000)
SFT
0.812
0.774
0.348
0.579
JPO
0.869
0.836
0.452
0.695
Multi-defendant ( n =800)
SFT
0.701
0.643
0.246
0.437
JPO
0.783
0.724
0.351
0.558
Appendix
Table 17: Generalization to complex cases (Qwen2.5-7B-Instruct, no retraining). The two harder subsets are drawn from held-out judgments outside JPO-Dataset.
Logic-weight granularity
Art. F1
Charge F1
Sent.
Full-Chain
Cost
Entropy-only (no logic weight)
0.923
0.912
0.523
0.769
1.00 ×
Stage-level (proposed)
0.929
0.921
0.536
0.791
1.05 ×
Span-level (per-sentence)
0.931
0.923
0.540
0.797
1.4 ×
Token-level (entailment model)
0.933
0.925
0.544
0.802
2.6 ×
Appendix
Table 18: Granularity of the legal-logic token weight (Qwen2.5-3B-Instruct, JPO-Dataset). Cost is training time relative to the entropy-only baseline.
Model
LLM-judge rating (1–5)
Win rate vs. SFT
SFT
3.14
–
Vanilla PPO
3.36
57%
Issue Tree Rubrics
3.53
63%
JPO
3.90
74%
Appendix
Table 19: Format-independent LLM-as-judge evaluation (Qwen2.5-3B-Instruct, 500 cases, judge = GPT-5.2). The judging rubric does not mention the four-step template.
Model
Fact
Article
Charge
Sentence
Pretrained
0.62
0.51
0.49
0.28
SFT
0.71
0.66
0.59
0.47
JPO
0.88
0.81
0.77
0.70
Appendix
Table 20: Within-subjects expert evaluation on Qwen2.5-7B-Instruct, averaged over 3 experts × 800 cases. Fleiss’ κ=0.73 .
Teacher prompt
SFT Sent.
JPO Sent.
JPO counterfact.
Unconstrained (answer-first)
0.404
0.522
0.60
Constrained forward (ours)
0.422
0.551
0.83
Appendix
Table 21: Teacher-prompt ablation (Qwen2.5-7B-Instruct, JPO-Dataset). The last column reports the counterfactual consistency rate of the resulting JPO model.
Table 23: From a generic RL recipe to JPO (Qwen2.5-3B-Instruct, JPO-Dataset). The legal composite reward combines the structure and consistency terms.
Method
Setting
Art. F1
Charge F1
Sent.
4-Step
Full-Chain
DeepSeek-V3.2
Zero-Shot
0.924
0.908
0.512
0.958
0.771
DeepSeek-V3.2
ICL (3-shot)
0.923
0.911
0.510
0.961
0.773
GPT-5.2
Zero-Shot
0.912
0.894
0.489
0.947
0.744
GPT-5.2
ICL (3-shot)
0.916
0.899
0.485
0.952
0.747
Qwen2.5-3B
JPO
0.929
0.921
0.536
0.967
0.791
Qwen2.5-3B
JPO + ICL (3-shot)
0.930
0.921
0.538
0.968
0.793
Appendix
Table 24: In-context learning on JPO-Dataset, extending Table 8 .
System
Params
Inference cost / 1k cases
One-time training
DeepSeek-V3.2 (API)
–
≈ $1.9
–
GPT-5.2 (API)
–
≈ $9.7
–
JPO Qwen2.5-3B (local)
3B
≈ $0.16
53 GPU-h
Appendix
Table 25: Cost comparison per 1,000 JPO-Dataset cases. API entries use public list pricing; the local entry amortizes A800 GPU time over inference.
Teacher
Scale
SFT Sent.
JPO Art.
JPO Charge
JPO Sent.
Qwen2.5-7B-Instruct
7B
0.401
0.930
0.919
0.532
Qwen2.5-32B-Instruct
32B
0.412
0.934
0.923
0.544
Qwen2.5-72B-Instruct (default)
72B
0.422
0.937
0.928
0.551
DeepSeek-V2
MoE
0.438
0.939
0.926
0.554
Claude Opus 4.8
prop.
0.449
0.941
0.930
0.558
Appendix
Table 26: Teacher-model scaling with a fixed student (Qwen2.5-7B-Instruct) on JPO-Dataset. “SFT Sent.” is the sentence score of the Stage-I checkpoint before RL.
Figure 6: Prompt template used to generate four-stage legal rationales for structured supervised fine-tuning.
Legal Judgment Prediction (LJP) has become a core benchmark for evaluating AI in the criminal legal domain, but it only sees criminal cases that have already passed prosecutorial review and been formally indicted. As a result, LJP leaves a substantial blind spot in assessing criminal liability, overlooking cases involving insufficient evidence, no criminal liability, or guilt exempted from punishment. To fill this gap, we propose \textbf{Prosecution Decision Prediction (PDP)}, the first Legal AI task built around prosecutorial review, which classifies each case into prosecution or one of three non-prosecution decisions and reflects legal AI's capabilities in evidence evaluation, legal subsumption, and value-based discretion. We further construct \textbf{PDP-Bench}, a benchmark of 4{,}630 real Chinese prosecutorial decisions spanning 190 charges. Extensive experiments show that state-of-the-art LLMs perform substantially worse on PDP than on LJP and that mainstream enhancement routes fail to close the gap. Moreover, controlled RLVR interventions show that simple outcome rewards fail to produce generalizable PDP discrimination.
Junyu Lu, Qi Wei, Peishuo Zheng +6
1Beijing Institute of Technology, Zhuhai · 2Osaka University · 3Xi’an University of Technology +2
Automating the drafting of judgment documents is pivotal to judicial efficiency, yet it remains challenging due to the dual requirements of comprehensive retrieval of legal information and rigorous logical reasoning. Existing approaches, typically relying on standard Retrieval-Augmented Generation and Supervised Fine-Tuning, often suffer from insufficient evidence recall, hallucinated statutory references, and logically flawed legal reasoning. To bridge this gap, we propose Judge-R1, a unified framework designed to enhance LLM-based judgment document generation by jointly improving legal information collection and judgment document generation. First, we introduce Agentic Legal Information Collection, which employs a dynamic planning agent to retrieve precise statutes and precedents from multiple sources. Second, we implement Rubric-Guided Optimization, a reinforcement learning phase utilizing Group Relative Policy Optimization (GRPO) with a comprehensive legal reward function to enforce adherence to judicial standards and reasoning logic. Extensive experiments on the JuDGE benchmark demonstrate that Judge-R1 significantly outperforms state-of-the-art baselines in both legal accuracy and generation quality.
Weihang Su, Xuanyi Chen, Yueyue Wu +2
Tsinghua University · Beijing, China · Quan Cheng Laboratory +1
What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=320), overall accuracy collapses from 0.500 (chance) to 0.072 (McNemar p < 10^-36), driven entirely by the rate of properly formatted answers falling from 0.900 to 0.109. The model stops committing to answers. Yet when it does commit, accuracy rises from 0.556 to 0.657, showing that the collapse is not a failure of capability but a strategic response: the model has learned that verbose responses packed with citations but empty of a direct answer score higher than terse correct ones. We term this the Saul Goodman effect, a policy that becomes maximally lawyerly while becoming maximally noncommittal, and prove formally that it is the optimal response to any surface feature proxy that attaches no penalty to abstention. We further show that 89.3% of citations produced after training are structurally implausible hallucinations, many of them subtly corrupted names of real landmark cases, constructed in effect to survive a casual read and fail under scrutiny. To detect this failure mode before deployment, we introduce three diagnostic tools: the Confidence Theater Score (CTS), the Citation Plausibility Rate (CPR), and the Regret Gap (RG). In a domain where a confidently wrong answer can constitute malpractice, the broader lesson is direct: a reward function that measures how legal a response looks will produce a model that is maximally photogenic and minimally useful.