cs.CLAug 30, 2026

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

Authors: Zhaolu Kang, Yantao Liu, Tailong Luo, Leqi Zheng, Lei Wei, Chenghua Zhu, Junhao Gong, Jiachen Qian, +9 more

Organizations: Hunyuan, Tencent · Peking University · Tsinghua University · City University of Hong Kong · University of California · University of Illinois Urbana-Champaign · Zhejiang University · University of Hong Kong

Abstract

Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, and while some have attempted to evaluate reasoning quality, their evaluations are indirect, often relying on LLM-generated rubrics that reflect model-internal preferences rather than the inherent logical structure of legal adjudication. We propose Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction. JPO first uses teacher-generated rationales to supervise a standardized four-step reasoning process, and then applies reinforcement learning with a composite reward over legal prediction quality, reasoning structure completeness, and cross-step consistency. JPO further introduces token-level advantage reweighting and adaptive clipping for legally salient reasoning segments. Experiments on multiple open-source language models and three Chinese legal benchmarks show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment

    May 27, 2026Junyu Lu, Qi Wei, Peishuo Zheng +6MljaildeAssessment

  2. Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization

    May 3, 2026Weihang Su, Xuanyi Chen, Yueyue Wu +2Legal Reasoning TasksReward Functions