Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation
Authors: Tobias Deußer, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa
Organizations: University of Bonn, Bonn, Germany · Lamarr-Institute for Machine Learning and Artificial Intelligence, Bonn, Germany · Universitat Rovira i Virgili, Reus, Spain · Fraunhofer IAIS, Sankt Augustin, Germany
Financial institutions operate under dense, frequently amended rulebooks, and answering a compliance question correctly requires not only fluency but verifiable grounding in the authoritative text. Large language models are attractive for this task, yet the models that firms can realistically deploy on-premise are compact ones, and compact models hallucinate obligations. We study whether a carefully domain-adapted retrieval-augmented generation pipeline closes that gap. Our retriever is built in three stages on top of LegalBERT: entailment tuning that recasts question--passage matching as premise--hypothesis reconstruction, contrastive tuning with in-batch negatives, and score-level fusion with BM25. Our generator is a compact model (2B--12B parameters) served under 4-bit quantization, either prompted or adapted with retrieval-aware fine-tuning (RAFT) through LoRA. On ObliQA, a question-answering benchmark built from the Abu Dhabi Global Market rulebooks, the staged retriever raises Recall@10 from 0.256 to 0.774 and outperforms BM25 (0.678) and E5-large-v2 (0.758), the strongest general-purpose dense encoder we tested. RAFT-LoRA then improves the composite RePASs answer-quality score for every model we could adapt, with the largest gain on the weakest one. However, the adapted models do not transfer to Australian case-law questions, and a closed-book model that receives no passages at all scores within 0.011 RePASs of the full pipeline while producing answers that cite nothing and misstate obligations. The retrieval gain is therefore measured directly, the generation gain is a gain in RePASs rather than demonstrated grounding, and grounding itself requires an evaluation protocol that RePASs does not provide.
Figures & tables
Dataset
Split
Documents
Questions
Corpus passages
ObliQA (ADGM)
train
32
20,573
13,705
validation
6
3,191
test
2
3,116
Open Australian
train
–
1,695
2,114
LegalQA
evaluation
–
427
Table 1: Datasets after leakage-free re-splitting. ObliQA splits are disjoint at document level; Open Australian LegalQA is grouped by legal citation.
Component
Setting
Base models
Qwen2.5-7B, Teuken-7B v0.6, Gemma-2-2B
Quantisation
4-bit NF4, FP16 compute
LoRA rank r / α / dropout
16 / 32 / 0.05
Target modules
q_proj , v_proj
Optimiser
paged AdamW (8-bit)
Learning rate / warm-up
2×10−4 / 0.05
Table 2: RAFT dataset construction and QLoRA adaptation settings, shared across all adapted models.
Retriever stage
R@10
MRR@10
nDCG@10
ObliQA (ADGM rulebooks)
LegalBERT (baseline)
0.256
0.158
0.181
+ entailment tuning (ET)
0.495
0.322
0.363
+ contrastive tuning (CT)
0.732
0.518
0.569
+ BM25 fusion (hybrid)
0.774
0.594
0.638
Open Australian LegalQA
Table 3: Staged retriever ablation. Each stage is initialized from the previous one; the hybrid adds BM25 at α=0.8 .
Retriever
Type
R@10
MRR@10
nDCG@10
BM25
lexical
0.678
0.504
0.546
BGE-M3
dense
0.711
0.544
0.584
BGE-base-en-v1.5
dense
0.719
0.549
0.590
E5-base-v2
dense
0.721
0.557
0.597
E5-large-v2
dense
0.758
0.595
0.635
Hybrid ET+CT (ours)
lexical+dense
0.774
0.594
0.638
Table 4: ObliQA retrieval against standard baselines, full validation split (3,733 question–passage pairs, 13,705-passage corpus).
Model
Strategy
Es↑
Cs↓
OCs↑
RePASs ↑
Qwen2.5-7B
zero-shot
0.894
0.189
0.301
0.668
few-shot
0.971
0.167
0.371
0.725
few-shot + CoT
0.957
0.154
0.359
0.721
RAFT-LoRA
0.926
0.157
0.409
0.725
Teuken-7B
zero-shot
0.944
0.256
0.332
0.673
v0.6
few-shot
0.928
0.200
0.259
0.662
Table 5: Answer generation on ObliQA, 150 evaluation questions, hybrid retriever context. Best RePASs per model in bold.
Model
Strategy
Es↑
Cs↓
OCs↑
RePASs ↑
Qwen2.5-7B
zero-shot
0.964
0.414
0.377
0.642
few-shot
0.964
0.402
0.350
0.637
few-shot + CoT
0.952
0.492
0.242
0.567
RAFT-LoRA †
0.957
0.451
0.238
0.581
Teuken-7B
zero-shot
0.943
0.552
0.288
0.559
v0.6
few-shot
0.941
0.578
0.212
0.525
Table 6: Answer generation on Open Australian LegalQA, 150 evaluation questions. † marks adapters trained on ObliQA and applied without retraining.
Model
Setting
Es↑
Cs↓
OCs↑
RePASs ↑
Qwen2.5-7B
closed-book
0.883
0.081
0.169
0.657
RAG (zero-shot)
0.894
0.189
0.301
0.668
Teuken-7B v0.6
closed-book
0.836
0.096
0.247
0.662
RAG (zero-shot)
0.944
0.256
0.332
0.673
Table 7: Closed-book control on ObliQA. Scores are computed against retrieved passages that the closed-book models did not receive. Closed-book runs cover 132 (Qwen) and 89 (Teuken) validation questions.
As large language models (LLMs) are increasingly deployed in financial services, a single non-compliant interaction can expose institutions to regulatory penalties and direct consumer harm. Existing guard models are built around general harm taxonomies and overlook violations grounded in specific financial regulations. We address this gap with a regulation-driven pipeline that operates directly on regulatory documents, inducing a financial compliance risk taxonomy and synthesizing grounded training data without any predefined violation categories. Instantiating the pipeline on Chinese financial regulations, we release \textbf{FinGuard-Bench}, to our knowledge the first benchmark for financial regulatory compliance detection, with expert-annotated labels at both the query and response levels. We further train \textbf{FinGuard}, a financial compliance detection model built on Qwen3-8B and trained on the regulation-grounded data via supervised fine-tuning and self-play reinforcement learning. On FinGuard-Bench, FinGuard substantially outperforms all baselines, including dedicated guard models and much larger general-purpose LLMs such as Qwen3.5-397B-A17B and GPT-5.1. Furthermore, FinGuard also preserves general safety capabilities and adapts to unseen institution-specific policies using policy documents alone. We will publicly release the code, prompts, and resources used in this work on GitHub.
Huaixia Dou, Jie Zhu, Minghao Wu +5
Qwen DianJin Team, Alibaba Cloud Computing · Tongyi Lab, Alibaba Group · School of Computer Science and Technology, Soochow University
Large language models (LLMs) are rapidly being adopted across various domains. However, their adoption in banking industry faces resistance due to demands for high accuracy, regulatory compliance, and the need for verifiable and grounded responses. We present a unified, data-efficient framework for training grounded domain-specific LLMs that optimizes answer quality, citation grounding, and calibrated refusal under real-world deployment constraints. First, we describe a data generation pipeline that combines LLM-as-a-Judge filtering, citation annotation, and curriculum learning with only 143M tokens. The resulting 12B model achieves high answer quality outperforming GPT-4.1 on citation grounding, with a modest citation tradeoff versus the untuned base. Second, we propose a calibrated refusal mechanism: training on 22% unanswerable examples yield a 12% "I don't know" rate, substantially improving over the base model's unsafe 4.3% rate while avoiding GPT-4.1's over-refusal (20.2%). Third, we present an end-to-end methodology spanning from data curation to quantized serving. The system is deployed at 40+ financial institutions, achieving a 7.1 percentage point improvement in query resolution (p < 0.001). Additionally, the model delivers 3-5x faster responses at 20-50x lower cost compared to GPT-4.1.
Denys Katerenchuk, Pablo Duboue, Keelan Evanini +6
Kasisto, New York, NY · Textualization, Vancouver, Canada · NBME, Philadelphia, PA
Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority structures. Unlike traditional multi-hop or legal QA, this task requires structured procedural lookups and evidence-set closure rather than entity resolution or case-law reasoning. Existing RAG systems struggle here due to flattened citation edges, fragmented retrieval expansions, and fragile post-hoc attribution. We formalize Regulatory Compliance QA with RegOps-Bench, a novel benchmark featuring an Operational Knowledge Graph derived from complex national R&D regulations. To address these bottlenecks, we propose RefWalk, a unified framework driven by a shared topic anchor. RefWalk traverses cross-document citations, fuses multi-view candidates via max-based aggregation, and enforces per-rule attribution to explicitly map claims to sources. We establish a strong baseline with substantial improvements in retrieval recall and citation accuracy. Finally, a contrastive evaluation on a U.S. health compliance dataset (HIPAA) reveals that existing systems exhibit saturation on flat-structure rules, underscoring the need for RegOps-Bench. Our code is available at https://github.com/yeongjoonJu/RefWalk.
Yeong-Joon Ju, Seong-Whan Lee
Department of Artificial Intelligence, Korea University