RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications
Authors: Jianpeng Zhao, Haihua Xu, Haoyang Zhang, Shuang Qian, Yixiang Tang, Xintao Wang, Kun Sun, Pei Wu, +2 more
Organizations: State Key Laboratory of Internet of Things for Smart City and Institute of Smart City Technologies, University of Macau, Macau SAR, China · ByteDance, Beijing, China
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.
Figures & tables
Figure 1: Answer correctness versus warrant completeness in RGDTs. Both traces reject correctly, but only CoT B states the waiver and coverage judgments. Red notes indicate omissions.
Figure 2: Warrant-access probes in RGDT-Bench. Circles denote rules ( r ), conditions ( c ), and evidence ( e ). Solid links connect rules to their conditions, and dashed links connect evidence to the conditions it supports. A rule may specify several conditions, and a condition may have multiple supporting passages or remain unresolved without support. P2/P3 illustrate evidence-dense and rule-dense contexts, with hard negatives in P2 and longer natural documents in P3. Support markings are illustrative and hidden from generators.
Figure 3: Response collection and diagnosis in RGDT-Bench. (a) A generator samples a native CoT and terminal decision from the RGDT input. (b) Using the outputs of (a), LLM-driven TWC extracts claims from the sampled CoT. WCV checks these claims and the terminal decision against the reference warrant, yielding condition-level diagnostics and a response-level completeness label.
Table 1: Correct-answer warrant gaps by generator (%). Final/Warr.: answer accuracy/warrant completeness. CWG′ : gap without the aggregation check. Equal-weight averages over supported task–probe combinations.
Figure 4: (a) Family-equal averages of Final (answer accuracy), Warrant (warrant completeness), and CWG (correct-answer warrant gap). (b) Each correct but incomplete response contributes equally to its failed layers. DS-R1 denotes DeepSeek-R1-Distill-Llama.
Table 3: Effects of trace input and supervision on RM assessment (%). Scores report the mean ± standard deviation across three training runs. Views include all responses or only correct answers. Outcome metrics are N/A in the correct-only view because it contains one outcome class. Bold/underline: best/second.
Figure 5: Warrant-supervision AUROC gains (pp). Cells show mean Warrant-RM minus Full-response ORM over three fits in both views. Gray: unsupported.
Warrant ↑
Accuracy ↑
N
Final
Full
WRM
Final
Full
WRM
All response groups
1
51.01
51.01
51.01
72.58
72.58
72.58
4
54.42
55.90
56.92
78.21
78.39
78.30
8
55.34
57.77
58.94
79.82
80.41
80.55
12
55.85
58.79
59.94
80.49
81.43
81.71
Table 4: Warrant-supervision effects on response selection (%). Means over three fits. Final/Full: Final-only/Full-response ORM; WRM: Warrant-RM. Bold/underline: best/second.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Layer
Required property
Representative failure
Rule scope
applied rules match a reference set where the comparison is determined, with unidentified applied rules marked NA
rule outside the reference set
Rule use
each condition judgment has its required rule support, using all required rule passages or an accepted alternative
missing required rule use
Condition
every required condition receives the reference judgment
omitted exception condition
Evidence
semantic support satisfies the source-specific requirement independently of literal citation format
evidence does not support the judgment
Aggregation
all required decision conditions enter aggregation, with their judgments combined according to the decision logic and consistent with the terminal decision
omitted aggregation input or inconsistent decision
Outcome
predicted label equals the final answer implied by the same reference warrant
final answer wrong
Appendix
Table 5: Warrant checks and separate diagnostics. The five completeness checks must pass. rule scope is a separate diagnostic. Failures shown are illustrative and may co-occur. C=1 denotes evaluability in Equation 2 .
Task track
Train
Valid.
Test
All
P1
P2
P3
Policy
56
20
18
94
✓
–
–
Contract
66
21
21
108
✓
✓
✓
Regulation
78
27
27
132
✓
–
–
Transaction
70
24
24
118
✓
✓
✓
Source cases
270
92
90
452
904 active prompts
Appendix
Table 6: Source-case counts by task and split. Each case yields one prompt per supported probe (checkmarks), with all probes sharing the case split. Native labels are balanced within each supported task–split–probe combination.
Statistic
Train
Validation
Test
All splits
Response Samples
Probe 1
25,920
8,832
8,640
43,392
Probe 2
13,056
4,320
4,320
21,696
Probe 3
13,056
4,320
4,320
21,696
Total
52,032
17,472
17,280
86,784
Condition Slots (K)
Appendix
Table 7: Data and supervision scale by probe. Tokens include the prompt, reasoning trace, and formatting text (Qwen3.5-9B tokenizer). K/M: thousands/millions. Totals are rounded after summation.
Collection
Train
Validation
Test
Intended use
Prompts
Balanced prompts
432
144
144
Balanced task–probe coverage, labels, and input lengths
Responses
Natural responses
52,032
17,472
17,280
Generated responses before balancing
Evaluable responses
51,832
17,400
17,203
Responses with evaluable warrant labels
Balanced response subset
11,856
3,720
3,564
Evaluable subset for training and evaluator comparison
Appendix
Table 8: Prompt and response counts by split. The first row counts prompts, and the remaining rows count responses. Sections 5.3 – 5.4 share the balanced evaluation responses.
Figure 6: Prompt–trace lengths in the balanced data. (a) Cumulative distributions by split. (b) Density by probe, with dark marks denoting medians. P1 covers four tasks, whereas P2/P3 cover Contract and Transaction .
Stage/model
Prompt
Native h
y^
TWC
Ref.
Output
Generation and Annotation
Generator
✓
×
×
×
×
native trace + terminal decision
TWC canonicalizer
✓
✓
×
×
×
explicit claims + trace locations
WCV verifier
×
×
✓
✓
✓
layer findings + warrant label
Evaluators
Existing evaluator
✓
✓
✓
×
×
warrant score
Appendix
Table 9: Information available to each model or stage. Checkmarks denote visible inputs. Ref.: reference warrant; TWC: extracted claims; h : reasoning trace; y^ : terminal decision. WCV verifies TWC claims using references withheld from generators and evaluators. Existing evaluators: Section 5.3 .
Task track
What must be judged?
What support is needed?
How is the decision checked?
Policy
The answer to the policy question
Use the supplied evidence or rules to justify the judgment without a preassigned passage.
Check that the stated grounds justify the answer.
Contract
Two to six condition judgments
For each condition, use the required evidence (or an accepted alternative) and its governing rule.
Combine all reference-required condition judgments into the decision.
Regulation
One to seven condition judgments
For each condition, use all jointly required evidence and rules, or an accepted alternative.
Apply the reference logic specifying which conditions must hold together or may serve as alternatives.
Transaction
Identify the transaction action and judge its overall compliance
Use supplied evidence to justify compliance and apply an accepted rule from the reference.
Check the grounds for compliance and agreement with the final decision. Check the transaction action separately.
Appendix
Table 10: What warrant checking requires for each task.
Group
n
α
Agreement
Pcomplete
Rcomplete
Transaction
108
0.654
83.3
–
–
Policy
36
0.894
97.2
–
–
Contract
162
0.635
83.3
–
–
Regulation
54
0.601
79.6
–
–
Overall
360
0.687
84.2
0.927
0.785
Appendix
Table 11: Automatic annotation versus human review. Agreement is reported as a percentage (%). Precision and recall are computed with complete warrants as the positive class.
Weights
Min–max
Top 1%
Top 5%
Kish ESS
ESS/ n
Shared
0.195–14.676
7.88%
19.84%
5,427.6
45.78%
Warrant-derived
0.283–3.646
3.44%
11.89%
9,996.6
84.32%
Appendix
Table 12: Training-weight concentration under the shared and warrant-derived schemes.
(a) Primary Attribution
Layer
Involv.
Single
Shared
Total
Rule use
44.5
31.8
4.6
36.4
Condition
32.3
2.8
12.4
15.2
Evidence
33.9
20.0
5.2
25.2
Aggregation
41.6
9.5
13.6
23.2
Appendix
Table 13: Failure attribution and weighting sensitivity (%). (a) Failure involvement and attribution shares. (b) Task-equal comparisons across all probes and common P1.
Table 14: Matched changes in warrant completeness (pp). Positive values favor P2/P3 over P1. Mean gives equal weight to the twelve displayed estimates from Contract and Transaction .
Table 15: Probability quality and warrant classification. Bold/underline: best/second.
Contrast
AUROC
AUPRC
BalAcc
Warrant Target
Full-response − Final-only
+5.68
+4.58
+2.96
Warrant − Full-response
+6.94
+10.16
+3.13
Warrant − Final-only
+12.62
+14.74
+6.09
Outcome Target
Full-response − Final-only
+1.92
+0.30
+0.88
Appendix
Table 16: Paired effects of trace access and supervision in the full view (pp). Differences compare metric means across fitted runs.
Table 17: Correct-only warrant AUROC by task (%). Tasks average supported probes, and RM rows average three fits. Mean: task-equal average. Bold/underline: best/second.
Scope
Full-response − Final-only
Warrant − Full-response
Transaction
+13.32
+5.08
Policy
+17.11
+13.84
Contract
−2.71
+17.37
Regulation
+0.34
+5.18
Appendix
Table 18: Correct-only AUROC gains from trace access and warrant supervision (pp). All models share training and validation weights.
Figure 7: Score-rank distributions among correct answers. Axes show weighted score percentiles averaged over three fits. Darker hexagons contain more within-panel response weight. Crosses mark median percentiles. Above the diagonal, Warrant-RM ranks responses higher than Full-response ORM, and below it, lower.
Metric and view
Balanced using warrant labels
Balanced using outcome labels (main)
AUROC ↑
Warrant completeness, full view
77.81±0.75
79.93±1.85
Warrant completeness, correct-only
69.32±0.65
69.24±2.15
Outcome correctness, full view
87.12±1.18
91.91±0.91
Best-of- N Selection ↑
Warrant completeness@ 16
70.13±0.98
71.52±1.30
Appendix
Table 19: Warrant-RM under two weighting schemes (%). Mean ± SD over three fits. Selection uses target-specific opportunity groups. Bold/underline: larger/smaller mean.
Figure 8: Selection gains. Bands show pointwise 95% source-group jackknife intervals. Vertical scales differ by row.
Warrant ↑
Accuracy ↑
N
Random
Oracle
Final
Full
WRM
Random
Oracle
Final
Full
WRM
All response groups
1
51.01
51.01
51.01
51.01
51.01
72.58
72.58
72.58
72.58
72.58
2
51.01
59.67
53.12
53.76
54.34
72.58
78.88
75.98
75.95
75.75
3
51.01
63.75
53.95
55.06
55.92
72.58
81.77
77.38
77.44
77.28
4
51.01
66.31
54.42
55.90
56.92
72.58
83.56
78.21
78.39
78.30
Appendix
Table 20: Complete selection results in both evaluation views (%). Means over three fits. Final/Full: Final-only/Full-response ORM; WRM: Warrant-RM. Bold/underline: best/second among Final, Full, and WRM. Random: uniform choice; Oracle: target-specific upper bound.
Figure 9: One Contract case across three probes.
Figure 10: Generation instructions and a recorded response.
Figure 11: Claim extraction and warrant verification.