JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction
Organizations: Hunyuan, Tencent · Peking University · Tsinghua University · City University of Hong Kong · University of California · University of Illinois Urbana-Champaign · Zhejiang University · University of Hong Kong
Abstract
Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, and while some have attempted to evaluate reasoning quality, their evaluations are indirect, often relying on LLM-generated rubrics that reflect model-internal preferences rather than the inherent logical structure of legal adjudication. We propose Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction. JPO first uses teacher-generated rationales to supervise a standardized four-step reasoning process, and then applies reinforcement learning with a composite reward over legal prediction quality, reasoning structure completeness, and cross-step consistency. JPO further introduces token-level advantage reweighting and adaptive clipping for legally salient reasoning segments. Experiments on multiple open-source language models and three Chinese legal benchmarks show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.
Figures & tables
| Method | Set. | JPO-Dataset | CAIL2018 | LawBench | ||||||||||||
| Art. | Charge | Sent. | 4-Step | Full | Art. | Charge | Sent. | 4-Step | Full | Art. | Charge | Sent. | 4-Step | Full | ||
| (F1) | (F1) | (Score) | (Comp.) | (Chain) | (F1) | (F1) | (Score) | (Comp.) | (Chain) | (F1) | (F1) | (Score) | (Comp.) | (Chain) | ||
| Open-Source Results | ||||||||||||||||
| Qwen3-4B-Instruct | Pre-trained | 0.521 | 0.468 | 0.174 | 0.315 | 0.208 | 0.485 | 0.442 | 0.158 | 0.291 | 0.187 | 0.463 | 0.419 | 0.142 | 0.264 | 0.175 |
| SFT | 0.884 | 0.858 | 0.405 | 0.902 | 0.652 | 0.856 | 0.831 | 0.381 | 0.882 | 0.627 | 0.834 | 0.809 | 0.364 | 0.857 | 0.598 | |
| JPO | 0.931 | 0.916 | 0.542 | 0.966 | 0.789 | 0.903 | 0.888 | 0.514 | 0.951 | 0.758 | 0.872 | 0.861 | 0.485 | 0.935 | 0.724 | |
| Dataset | SFT Train | RL Train | Test |
| JPO-Dataset | 239,515 | 9,691 | 20,396 |
| CAIL2018 | – | – | 30,000 |
| LawBench | – | – | 1,500 |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Statistic | Value |
| Time range of collected cases | 2024–2026 |
| Number of unique charges | 192 |
| Number of unique statutory articles | 176 |
| Average fact length (tokens) | 217.1 |
| Median fact length (tokens) | 194 |
| Average number of articles per case | 1.04 |
| Method | Art. F1 | Charge F1 | Sent. | Full-Chain |
| JPO | 0.929 | 0.921 | 0.536 | 0.791 |
| w/o | 0.923 | 0.915 | 0.520 | 0.768 |
| w/o | 0.920 | 0.908 | 0.513 | 0.755 |
| w/o | 0.925 | 0.916 | 0.497 | 0.763 |
| Entropy-only weighting | 0.923 | 0.912 | 0.523 | 0.769 |
| Logic-only weighting | 0.924 | 0.914 | 0.525 | 0.771 |
| Art. F1 | Charge F1 | Sent. | |||
| 0.85 | 0.05 | 0.10 | 0.923 | 0.913 | 0.518 |
| 0.75 | 0.0625 | 0.1875 | 0.929 | 0.921 | 0.536 |
| 0.65 | 0.0625 | 0.2875 | 0.926 | 0.917 | 0.530 |
| 0.33 | 0.33 | 0.33 | 0.904 | 0.893 | 0.494 |
| Art. F1 | Charge F1 | Sent. | |
| 0.0 | 0.922 | 0.911 | 0.519 |
| 0.3 | 0.924 | 0.915 | 0.523 |
| 0.6 | 0.929 | 0.921 | 0.536 |
| 0.9 | 0.927 | 0.918 | 0.531 |
| Art. F1 | Charge F1 | Sent. | |
| 0.0 | 0.924 | 0.914 | 0.525 |
| 0.25 | 0.926 | 0.917 | 0.529 |
| 0.5 | 0.929 | 0.921 | 0.536 |
| 0.75 | 0.925 | 0.916 | 0.527 |
| 1.0 | 0.923 | 0.912 | 0.523 |
| Method | Set. | JPO-Dataset | CAIL2018 | LawBench | ||||||||||||
| Art. | Charge | Sent. | 4-Step | Full | Art. | Charge | Sent. | 4-Step | Full | Art. | Charge | Sent. | 4-Step | Full | ||
| (F1) | (F1) | (Score) | (Comp.) | (Chain) | (F1) | (F1) | (Score) | (Comp.) | (Chain) | (F1) | (F1) | (Score) | (Comp.) | (Chain) | ||
| DeepSeek-V3.2 | Zero-Shot | 0.924 | 0.908 | 0.512 | 0.958 | 0.771 | 0.896 | 0.875 | 0.481 | 0.942 | 0.738 | 0.864 | 0.841 | 0.439 | 0.925 | 0.706 |
| Qwen3-32B | Zero-Shot | 0.917 | 0.911 | 0.504 | 0.961 | 0.759 | 0.887 | 0.882 | 0.473 | 0.945 | 0.722 | 0.858 | 0.835 | 0.428 | 0.931 | 0.694 |
| GPT-5.2 | Zero-Shot | 0.912 | 0.894 | 0.489 | 0.947 | 0.744 | 0.879 | 0.861 | 0.455 | 0.933 | 0.705 | 0.845 | 0.829 | 0.412 | 0.916 | 0.683 |
| Claude-Sonnet-4.5 | Zero-Shot | 0.908 | 0.891 | 0.493 | 0.952 | 0.751 | 0.882 | 0.866 | 0.462 | 0.937 | 0.713 | 0.849 | 0.824 | 0.417 | 0.918 | 0.687 |
| Method | Paradigm | Art. F1 | Charge F1 | Sentencing (as reported) |
| TopJudge | dependency-aware | 0.737 | 0.800 | 0.357 (Acc) |
| MPBFN | dependency-aware | 0.706 | 0.757 | 0.362 (Acc) |
| LADAN | structure & knowledge | 0.738 | 0.801 | 0.361 (Acc) |
| NeurJudge | structure & knowledge | 0.797 | 0.807 | 0.374 (Acc) |
| CTM | structure & knowledge | 0.768 | 0.780 | 0.374 (Acc) |
| EPM | dependency-aware | 0.781 | 0.814 | 0.367 (Acc) |
| System | Fact | Article | Charge | Sentence |
| SFT (Qwen2.5-7B-Instruct) | 0.73 | 0.63 | 0.58 | 0.49 |
| JPO (Qwen2.5-7B-Instruct) | 0.87 | 0.83 | 0.79 | 0.71 |
| Improvement | +0.14 | +0.20 | +0.21 | +0.22 |
| Teacher | Stage | Art. (F1) | Charge (F1) | Sent. (Score) |
| Qwen2.5-72B-Instruct | SFT | 0.893 | 0.871 | 0.422 |
| JPO | 0.937 | 0.928 | 0.551 | |
| DeepSeek-V2 | SFT | 0.901 | 0.875 | 0.438 |
| JPO | 0.939 | 0.926 | 0.554 |
| Model | Pre-trained | SFT | JPO | |
| Qwen2.5-3B-Instruct | 1.0 | 0.317 | 0.613 | 0.747 |
| 2.0 | 0.185 | 0.472 | 0.614 | |
| 3.0 | 0.114 | 0.391 | 0.536 | |
| 4.0 | 0.086 | 0.347 | 0.475 | |
| 5.0 | 0.067 | 0.305 | 0.437 | |
| Qwen2.5-7B-Instruct | 1.0 | 0.430 | 0.657 | 0.765 |
| Method | GPU Hours | Avg. Length | Std. Full-Chain |
| SFT | 34 | 456 | 0.018 |
| Vanilla PPO | 49 | 471 | 0.025 |
| JPO | 53 | 478 | 0.015 |
| Error Type | Proportion |
| Missing secondary article | 27.5% |
| Correct charge but wrong sentence band | 25.3% |
| Confusion between neighboring charges | 21.6% |
| Incomplete mitigation/aggravation reasoning | 16.7% |
| Formatting or parsing errors | 8.9% |
| Automatic proxy | Expert step judged | Spearman |
| (fact-to-article) | Statutory analysis | 0.67 |
| (article-to-charge) | Charge determination | 0.71 |
| (charge-to-sentence) | Sentence prediction | 0.64 |
| Full-Chain Consistency | Overall reasoning | 0.72 |
| Reasoning structure | Art. F1 | Charge F1 | Sent. | Full-Chain |
| Standard CoT (free-form) | 0.921 | 0.907 | 0.503 | 0.724 |
| Fact–Element–Charge (3-step) | 0.928 | 0.914 | 0.517 | 0.749 |
| Issue-tree decomposition | 0.931 | 0.918 | 0.525 | 0.768 |
| Four-step F A C S | 0.937 | 0.928 | 0.551 | 0.806 |
| Evaluation subset | Model | Art. F1 | Charge F1 | Sent. | Full-Chain |
| Single-def./single-charge | SFT | 0.893 | 0.871 | 0.422 | 0.681 |
| JPO | 0.937 | 0.928 | 0.551 | 0.806 | |
| Multi-charge ( =1,000) | SFT | 0.812 | 0.774 | 0.348 | 0.579 |
| JPO | 0.869 | 0.836 | 0.452 | 0.695 | |
| Multi-defendant ( =800) | SFT | 0.701 | 0.643 | 0.246 | 0.437 |
| JPO | 0.783 | 0.724 | 0.351 | 0.558 |
| Logic-weight granularity | Art. F1 | Charge F1 | Sent. | Full-Chain | Cost |
| Entropy-only (no logic weight) | 0.923 | 0.912 | 0.523 | 0.769 | 1.00 |
| Stage-level (proposed) | 0.929 | 0.921 | 0.536 | 0.791 | 1.05 |
| Span-level (per-sentence) | 0.931 | 0.923 | 0.540 | 0.797 | 1.4 |
| Token-level (entailment model) | 0.933 | 0.925 | 0.544 | 0.802 | 2.6 |
| Model | LLM-judge rating (1–5) | Win rate vs. SFT |
| SFT | 3.14 | – |
| Vanilla PPO | 3.36 | 57% |
| Issue Tree Rubrics | 3.53 | 63% |
| JPO | 3.90 | 74% |
| Model | Fact | Article | Charge | Sentence |
| Pretrained | 0.62 | 0.51 | 0.49 | 0.28 |
| SFT | 0.71 | 0.66 | 0.59 | 0.47 |
| JPO | 0.88 | 0.81 | 0.77 | 0.70 |
| Teacher prompt | SFT Sent. | JPO Sent. | JPO counterfact. |
| Unconstrained (answer-first) | 0.404 | 0.522 | 0.60 |
| Constrained forward (ours) | 0.422 | 0.551 | 0.83 |
| Model | Counterfactual consistency rate |
| Pretrained | 0.43 |
| SFT | 0.61 |
| JPO | 0.83 |
| Configuration | Art. F1 | Charge F1 | Sent. | Full-Chain |
| Structured four-step SFT | 0.873 | 0.847 | 0.391 | 0.622 |
| + generic PPO (outcome-only reward) | 0.896 | 0.869 | 0.454 | 0.678 |
| + generic entropy-only token weighting | 0.900 | 0.875 | 0.466 | 0.690 |
| + legal composite reward | 0.919 | 0.908 | 0.501 | 0.747 |
| JPO (+ legal-logic token opt.) | 0.929 | 0.921 | 0.536 | 0.791 |
| Method | Setting | Art. F1 | Charge F1 | Sent. | 4-Step | Full-Chain |
| DeepSeek-V3.2 | Zero-Shot | 0.924 | 0.908 | 0.512 | 0.958 | 0.771 |
| DeepSeek-V3.2 | ICL (3-shot) | 0.923 | 0.911 | 0.510 | 0.961 | 0.773 |
| GPT-5.2 | Zero-Shot | 0.912 | 0.894 | 0.489 | 0.947 | 0.744 |
| GPT-5.2 | ICL (3-shot) | 0.916 | 0.899 | 0.485 | 0.952 | 0.747 |
| Qwen2.5-3B | JPO | 0.929 | 0.921 | 0.536 | 0.967 | 0.791 |
| Qwen2.5-3B | JPO + ICL (3-shot) | 0.930 | 0.921 | 0.538 | 0.968 | 0.793 |
| System | Params | Inference cost / 1k cases | One-time training |
| DeepSeek-V3.2 (API) | – | $1.9 | – |
| GPT-5.2 (API) | – | $9.7 | – |
| JPO Qwen2.5-3B (local) | 3B | $0.16 | 53 GPU-h |
| Teacher | Scale | SFT Sent. | JPO Art. | JPO Charge | JPO Sent. |
| Qwen2.5-7B-Instruct | 7B | 0.401 | 0.930 | 0.919 | 0.532 |
| Qwen2.5-32B-Instruct | 32B | 0.412 | 0.934 | 0.923 | 0.544 |
| Qwen2.5-72B-Instruct (default) | 72B | 0.422 | 0.937 | 0.928 | 0.551 |
| DeepSeek-V2 | MoE | 0.438 | 0.939 | 0.926 | 0.554 |
| Claude Opus 4.8 | prop. | 0.449 | 0.941 | 0.930 | 0.558 |