Organizations: School of Computer Science and Technology, Fudan University · School of Integrated Circuits, Nanjing University · School of Economics, Fudan University
Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.
Figures & tables
Figure 1: Overview of InsClaimBench. An example illustrating the end-to-end claim adjudication process, including atomic rule judgments, adjudication modules, claim decisions, and payout outcomes under controlled factual changes.
Figure 2: Construction pipeline of InsClaimBench. Expert-defined rules and logic guide scenario generation, reference answer construction, and independent evaluation at the rule, module, and claim levels.
Figure 3: Rule extraction and organization in InsClaimBench. Left , claims experts and LLMs extract atomic rules from policy provisions, laws and regulations, and expert knowledge. Center , the reference decision structure links case evidence and atomic judgments to adjudication modules, claim outcomes, and payout amounts. Right , rules are organized into general, category-specific, and product-specific levels.
Insurance domain
Distinct rules
Total rule judgments
Cases
Motor
86
32,054
1,265
Property
28
35,000
1,250
Health
Critical illness
43
8,701
627
Medical
48
10,901
638
Total
—
86,656
3,780 (375)
Table 1: Rules, rule judgments, and cases across insurance domains in InsClaimBench.
Rule
Module
Claim output
Acc.
Vector EM
CCE
EC
CE
T/F Acc.
Amount Acc.
MAPE
93.99
92.46
93.47
97.61
95.13
96.79
96.32
2.61
Table 2: Human evaluation on the sampled subset (%). Each metric is computed separately for each of the four experts on their assigned subset and then averaged across experts.
Payout decision accuracy
Payout calculation
Model
Overall
Auto
Property
Illness
Medical
MAPE
Joint Acc.
GPT-5.6-Sol
80.05
82.61
79.52
82.62
73.51
5.63
69.72
Gemini-3.8-Flash
79.29
76.36
84.24
84.53
70.22
5.82
73.15
DeepSeek-V4.1-Flash
77.46
79.84
74.72
83.89
71.79
7.67
69.21
Qwen3.8-Flash
76.64
80.79
73.52
79.43
71.79
7.73
64.88
Kimi-K2.6
80.19
86.80
78.40
80.86
69.91
10.45
64.26
Table 3: Payout decision and amount performance (%). Decision accuracy is shown overall and by insurance type. Amount metrics are pooled across insurance types.
Model
Rule Acc.
Vector EM
R+D−
R−D+
GPT-5.6-Sol
95.48
36.90
5.05
48.20
Gemini-3.8-Flash
94.58
31.96
3.70
51.03
DeepSeek-V4.1-Flash
87.13
26.51
3.68
54.63
Qwen3.8-Flash
91.85
26.43
3.94
54.15
Kimi-K2.6
89.31
14.68
1.98
67.49
Qwen3.5-27B
91.48
15.03
2.65
61.85
Table 4: Rule accuracy and agreement with final T/F decisions (%). Metrics follow Section 4.3 .
Model
CCE
EC
CE
Acc.
Final fail. ( n )
Acc.
Final fail. ( n )
Acc.
Final fail. ( n )
GPT-5.6-Sol
92.57
62.63 (281)
73.39
42.25 (1,006)
91.67
59.05 (315)
Gemini-3.8-Flash
86.06
51.42 (527)
75.58
41.60 (923)
92.22
48.98 (294)
DeepSeek-V4.1-Flash
78.68
40.94 (806)
70.21
29.13 (1,126)
83.57
42.83 (621)
Qwen3.8-Flash
83.36
48.97 (629)
73.49
38.62 (1,002)
87.49
53.70 (473)
Kimi-K2.6
80.74
34.07 (728)
71.14
27.68 (1,091)
88.07
52.55 (451)
Table 5: Module accuracy and final T/F failure when a module is wrong (%). Error-case counts are in parentheses. Metrics follow Section 4.3 .
Model
Rule
Module
Payout outcomes
Sens.
Inv.
PCΔ
PC=
Sens.
Inv.
PCΔ
PC=
MUF
PAF ( n )
MME ( n )
GPT-5.6-Sol
84.80
96.40
84.22
94.32
67.79
89.68
67.71
84.67
29.55
11.25 (1,813)
34.23 (111)
Gemini-3.8-Flash
84.04
97.06
83.75
93.71
69.08
90.99
69.08
84.10
28.57
9.27 (2,027)
27.41 (135)
DeepSeek-V4.1-Flash
71.48
92.86
70.68
84.53
61.76
83.77
60.76
73.13
34.05
8.69 (1,415)
68.03 (147)
Qwen3.8-Flash
76.80
94.27
76.14
89.84
60.38
88.60
59.85
79.93
36.00
10.81 (1,453)
66.96 (112)
Kimi-K2.6
71.93
94.00
71.38
87.30
59.69
88.78
59.54
78.87
37.00
18.89 (1,212)
49.63 (135)
Table 6: Counterfactual responses and payout outcomes on case pairs (%). Eligible pair counts for PAF and MME are in parentheses. Metrics follow Section 4.3 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Expert
T
F
A
3,212
2,204
B
3,418
1,998
C
3,300
1,791
D
3,477
1,614
Appendix
Table 7: Distribution of expert judgments in human validation.
Group 1
Group 2
A:T
A:F
Agreement(%)
κ
C:T
C:F
Agreement(%)
κ
B:T
3,086
332
91.54
0.822
D:T
3,167
310
91.30
0.805
B:F
126
1,872
D:F
133
1,481
Appendix
Table 8: Inter-expert agreement in blinded expert quality verification.
Model
Top rule (error rate)
All cases
R−D+
GPT-5.6-Sol
C07 (44.00)
C07 (43.91)
Gemini-3.8-Flash
C07 (39.36)
P07 (44.89)
DeepSeek-V4.1-Flash
P01 (76.56)
P01 (77.11)
Qwen3.8-Flash
P01 (75.84)
P01 (75.25)
Kimi-K2.6
P21 (63.52)
P21 (68.00)
Appendix
Table 9: Frequent rule errors (%). Top rules are ranked by error count within each subset. All listed rules are from property insurance. Metrics follow Section 4.3 .
Insurance type
GPT-5.6 Sol
Gemini-3.8 Flash
DeepSeek Flash
Qwen3.8 Flash
Kimi-K2.6
Qwen3.5-27B
Auto
+0.00 [-1.11, +1.03]
-0.95 [-2.06, +0.16]
-0.79 [-1.90, +0.24]
+0.79 [-0.24, +1.90]
+0.16 [-0.95, +1.26]
-0.47 [-1.66, +0.71]
Property
+0.48 [-0.50, +0.75]
-0.24 [-0.78, +0.20]
-0.16 [-0.59, +0.09]
-0.08 [-0.57, +0.42]
-0.24 [-0.88, +0.81]
+0.24 [-0.33, +0.66]
Health
Illness
+0.77 [-0.67, +2.23]
+0.14 [-1.28, +1.44]
+1.09 [-0.64, +2.85]
+1.09 [-0.48, +2.68]
-0.51 [-2.07, +0.93]
-0.52 [-2.07, +0.92]
Medical
-0.41 [-1.50, +0.69]
-0.12 [-1.23, +0.98]
+0.21 [-0.88, +1.30]
-0.43 [-1.52, +0.67]
-0.44 [-1.53, +0.65]
-0.17 [-1.27, +0.94]
Appendix
Table 10: Case-generator ablation. Entries are final T/F accuracy differences (DeepSeek minus GPT, percentage points), with paired family-bootstrap 95% confidence intervals.
Rule
Proposition represented by T
r1
The claimant and policy satisfy the applicable eligibility requirements, and the relevant coverage is in force.
r2
The event falls within the insurance period or an applicable contractual extension.
r3
A valid disclosure-related bar to coverage is established.
r4
A medical event covered by the relevant benefit has occurred.
r5
The claimed medical expenses are related to the event, reasonable, and necessary.
r6
The waiting-period provision applies and the event falls within the waiting period.
Appendix
Table 11: Atomic judgments used in the illustrative medical-insurance example.