Organizations: School of Computer Science and Technology, Fudan University · School of Integrated Circuits, Nanjing University · School of Economics, Fudan University
Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.
Figures & tables
Figure 1: Overview of InsClaimBench. An example illustrating the end-to-end claim adjudication process, including atomic rule judgments, adjudication modules, claim decisions, and payout outcomes under controlled factual changes.
Figure 2: Construction pipeline of InsClaimBench. Expert-defined rules and logic guide scenario generation, reference answer construction, and independent evaluation at the rule, module, and claim levels.
Figure 3: Rule extraction and organization in InsClaimBench. Left , claims experts and LLMs extract atomic rules from policy provisions, laws and regulations, and expert knowledge. Center , the reference decision structure links case evidence and atomic judgments to adjudication modules, claim outcomes, and payout amounts. Right , rules are organized into general, category-specific, and product-specific levels.
Insurance domain
Distinct rules
Total rule judgments
Cases
Motor
86
32,054
1,265
Property
28
35,000
1,250
Health
Critical illness
43
8,701
627
Medical
48
10,901
638
Total
—
86,656
3,780 (375)
Table 1: Rules, rule judgments, and cases across insurance domains in InsClaimBench.
Rule
Module
Claim output
Acc.
Vector EM
CCE
EC
CE
T/F Acc.
Amount Acc.
MAPE
93.99
92.46
93.47
97.61
95.13
96.79
96.32
2.61
Table 2: Human evaluation on the sampled subset (%). Each metric is computed separately for each of the four experts on their assigned subset and then averaged across experts.
Payout decision accuracy
Payout calculation
Model
Overall
Auto
Property
Illness
Medical
MAPE
Joint Acc.
GPT-5.6-Sol
80.05
82.61
79.52
82.62
73.51
5.63
69.72
Gemini-3.8-Flash
79.29
76.36
84.24
84.53
70.22
5.82
73.15
DeepSeek-V4.1-Flash
77.46
79.84
74.72
83.89
71.79
7.67
69.21
Qwen3.8-Flash
76.64
80.79
73.52
79.43
71.79
7.73
64.88
Kimi-K2.6
80.19
86.80
78.40
80.86
69.91
10.45
64.26
Table 3: Payout decision and amount performance (%). Decision accuracy is shown overall and by insurance type. Amount metrics are pooled across insurance types.
Model
Rule Acc.
Vector EM
R+D−
R−D+
GPT-5.6-Sol
95.48
36.90
5.05
48.20
Gemini-3.8-Flash
94.58
31.96
3.70
51.03
DeepSeek-V4.1-Flash
87.13
26.51
3.68
54.63
Qwen3.8-Flash
91.85
26.43
3.94
54.15
Kimi-K2.6
89.31
14.68
1.98
67.49
Qwen3.5-27B
91.48
15.03
2.65
61.85
Table 4: Rule accuracy and agreement with final T/F decisions (%). Metrics follow Section 4.3 .
Model
CCE
EC
CE
Acc.
Final fail. ( n )
Acc.
Final fail. ( n )
Acc.
Final fail. ( n )
GPT-5.6-Sol
92.57
62.63 (281)
73.39
42.25 (1,006)
91.67
59.05 (315)
Gemini-3.8-Flash
86.06
51.42 (527)
75.58
41.60 (923)
92.22
48.98 (294)
DeepSeek-V4.1-Flash
78.68
40.94 (806)
70.21
29.13 (1,126)
83.57
42.83 (621)
Qwen3.8-Flash
83.36
48.97 (629)
73.49
38.62 (1,002)
87.49
53.70 (473)
Kimi-K2.6
80.74
34.07 (728)
71.14
27.68 (1,091)
88.07
52.55 (451)
Table 5: Module accuracy and final T/F failure when a module is wrong (%). Error-case counts are in parentheses. Metrics follow Section 4.3 .
Model
Rule
Module
Payout outcomes
Sens.
Inv.
PCΔ
PC=
Sens.
Inv.
PCΔ
PC=
MUF
PAF ( n )
MME ( n )
GPT-5.6-Sol
84.80
96.40
84.22
94.32
67.79
89.68
67.71
84.67
29.55
11.25 (1,813)
34.23 (111)
Gemini-3.8-Flash
84.04
97.06
83.75
93.71
69.08
90.99
69.08
84.10
28.57
9.27 (2,027)
27.41 (135)
DeepSeek-V4.1-Flash
71.48
92.86
70.68
84.53
61.76
83.77
60.76
73.13
34.05
8.69 (1,415)
68.03 (147)
Qwen3.8-Flash
76.80
94.27
76.14
89.84
60.38
88.60
59.85
79.93
36.00
10.81 (1,453)
66.96 (112)
Kimi-K2.6
71.93
94.00
71.38
87.30
59.69
88.78
59.54
78.87
37.00
18.89 (1,212)
49.63 (135)
Table 6: Counterfactual responses and payout outcomes on case pairs (%). Eligible pair counts for PAF and MME are in parentheses. Metrics follow Section 4.3 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Expert
T
F
A
3,212
2,204
B
3,418
1,998
C
3,300
1,791
D
3,477
1,614
Appendix
Table 7: Distribution of expert judgments in human validation.
Group 1
Group 2
A:T
A:F
Agreement(%)
κ
C:T
C:F
Agreement(%)
κ
B:T
3,086
332
91.54
0.822
D:T
3,167
310
91.30
0.805
B:F
126
1,872
D:F
133
1,481
Appendix
Table 8: Inter-expert agreement in blinded expert quality verification.
Model
Top rule (error rate)
All cases
R−D+
GPT-5.6-Sol
C07 (44.00)
C07 (43.91)
Gemini-3.8-Flash
C07 (39.36)
P07 (44.89)
DeepSeek-V4.1-Flash
P01 (76.56)
P01 (77.11)
Qwen3.8-Flash
P01 (75.84)
P01 (75.25)
Kimi-K2.6
P21 (63.52)
P21 (68.00)
Appendix
Table 9: Frequent rule errors (%). Top rules are ranked by error count within each subset. All listed rules are from property insurance. Metrics follow Section 4.3 .
Insurance type
GPT-5.6 Sol
Gemini-3.8 Flash
DeepSeek Flash
Qwen3.8 Flash
Kimi-K2.6
Qwen3.5-27B
Auto
+0.00 [-1.11, +1.03]
-0.95 [-2.06, +0.16]
-0.79 [-1.90, +0.24]
+0.79 [-0.24, +1.90]
+0.16 [-0.95, +1.26]
-0.47 [-1.66, +0.71]
Property
+0.48 [-0.50, +0.75]
-0.24 [-0.78, +0.20]
-0.16 [-0.59, +0.09]
-0.08 [-0.57, +0.42]
-0.24 [-0.88, +0.81]
+0.24 [-0.33, +0.66]
Health
Illness
+0.77 [-0.67, +2.23]
+0.14 [-1.28, +1.44]
+1.09 [-0.64, +2.85]
+1.09 [-0.48, +2.68]
-0.51 [-2.07, +0.93]
-0.52 [-2.07, +0.92]
Medical
-0.41 [-1.50, +0.69]
-0.12 [-1.23, +0.98]
+0.21 [-0.88, +1.30]
-0.43 [-1.52, +0.67]
-0.44 [-1.53, +0.65]
-0.17 [-1.27, +0.94]
Appendix
Table 10: Case-generator ablation. Entries are final T/F accuracy differences (DeepSeek minus GPT, percentage points), with paired family-bootstrap 95% confidence intervals.
Rule
Proposition represented by T
r1
The claimant and policy satisfy the applicable eligibility requirements, and the relevant coverage is in force.
r2
The event falls within the insurance period or an applicable contractual extension.
r3
A valid disclosure-related bar to coverage is established.
r4
A medical event covered by the relevant benefit has occurred.
r5
The claimed medical expenses are related to the event, reasonable, and necessary.
r6
The waiting-period provision applies and the event falls within the waiting period.
Appendix
Table 11: Atomic judgments used in the illustrative medical-insurance example.
Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, context-grounded, and tool-executable decisions. We introduce \textbf{INS-ActBench}, a comprehensive benchmark for evaluating professional actuarial capability in LLMs. INS-ActBench contains 12,050 Q&A pairs from public exams and sample questions released by 16 actuarial associations. It covers three subsets: \textbf{INS-Act-Know} for standardized actuarial knowledge, \textbf{INS-Act-Case} for long-context insurance case reasoning, and \textbf{INS-Act-Practice} for spreadsheet and R-code tasks with verifiable numerical outputs. Experiments on nine representative LLMs and human actuarial experts reveal a clear capability boundary: frontier LLMs perform strongly on standardized knowledge, but remain much weaker in case reasoning, tool-based workflows, and jurisdiction-sensitive practice. INS-ActBench provides a reproducible foundation for developing actuarial LLMs toward reliable professional assistance. The code is available at https://github.com/FDU-INS/INS-ActBench.
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Using parameter-efficient fine-tuning, small Qwen3 models (0.6B\unicodex20138B) reach 98.7% accuracy on label-flipping perturbations, outperforming larger zero-shot models and frontier systems (GPT-5.4 Pro (74.0%) and Gemini 2.5 Flash (73.9%)). The gains generalise to unseen perturbation types, indicating robust numerical decision boundaries rather than memorised edits. Robustness also transfers without target-domain data, significantly improving cross-lingual performance in Spanish. We further show that the same fine-tuning recipe confers robustness to evidence-side perturbations, using the VitaminC dataset.
Peter Røysland Aarnes, Vinay Setty
University of Stavanger · Factiverse AI and University of Stavanger
Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings remain unclear. Through controlled comparisons between reasoning and non-reasoning judges, we show that explicit reasoning substantially improves judgment accuracy on tasks requiring structured verification (e.g., math and coding), while offering limited or even negative gains on simpler evaluations and incurring significantly higher computational cost. These findings motivate that reasoning should be used selectively rather than universally, with awareness of possible distribution shift. We propose a Robust Adaptive Cost-Efficient Routing (RACER), which dynamically selects between reasoning and non-reasoning judges under a fixed budget by formulating routing as a constrained distributionally robust optimization problem. RACER explicitly accounts for distribution shift via a KL-divergence uncertainty set, admits an efficient primal--dual algorithm, and enjoys theoretical guarantees including uniqueness of the optimal policy and linear convergence. Extensive experiments show that RACER achieves superior accuracy--cost trade-offs under distribution shift.
Wenbo Zhang, Lijinghua Zhang, Liner Xiang +1
Department of Statistics, University of California, Irvine, USA.