Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation
Authors: Tobias Deußer, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa
Organizations: University of Bonn, Bonn, Germany · Lamarr-Institute for Machine Learning and Artificial Intelligence, Bonn, Germany · Universitat Rovira i Virgili, Reus, Spain · Fraunhofer IAIS, Sankt Augustin, Germany
Financial institutions operate under dense, frequently amended rulebooks, and answering a compliance question correctly requires not only fluency but verifiable grounding in the authoritative text. Large language models are attractive for this task, yet the models that firms can realistically deploy on-premise are compact ones, and compact models hallucinate obligations. We study whether a carefully domain-adapted retrieval-augmented generation pipeline closes that gap. Our retriever is built in three stages on top of LegalBERT: entailment tuning that recasts question--passage matching as premise--hypothesis reconstruction, contrastive tuning with in-batch negatives, and score-level fusion with BM25. Our generator is a compact model (2B--12B parameters) served under 4-bit quantization, either prompted or adapted with retrieval-aware fine-tuning (RAFT) through LoRA. On ObliQA, a question-answering benchmark built from the Abu Dhabi Global Market rulebooks, the staged retriever raises Recall@10 from 0.256 to 0.774 and outperforms BM25 (0.678) and E5-large-v2 (0.758), the strongest general-purpose dense encoder we tested. RAFT-LoRA then improves the composite RePASs answer-quality score for every model we could adapt, with the largest gain on the weakest one. However, the adapted models do not transfer to Australian case-law questions, and a closed-book model that receives no passages at all scores within 0.011 RePASs of the full pipeline while producing answers that cite nothing and misstate obligations. The retrieval gain is therefore measured directly, the generation gain is a gain in RePASs rather than demonstrated grounding, and grounding itself requires an evaluation protocol that RePASs does not provide.
Figures & tables
Dataset
Split
Documents
Questions
Corpus passages
ObliQA (ADGM)
train
32
20,573
13,705
validation
6
3,191
test
2
3,116
Open Australian
train
–
1,695
2,114
LegalQA
evaluation
–
427
Table 1: Datasets after leakage-free re-splitting. ObliQA splits are disjoint at document level; Open Australian LegalQA is grouped by legal citation.
Component
Setting
Base models
Qwen2.5-7B, Teuken-7B v0.6, Gemma-2-2B
Quantisation
4-bit NF4, FP16 compute
LoRA rank r / α / dropout
16 / 32 / 0.05
Target modules
q_proj , v_proj
Optimiser
paged AdamW (8-bit)
Learning rate / warm-up
2×10−4 / 0.05
Table 2: RAFT dataset construction and QLoRA adaptation settings, shared across all adapted models.
Retriever stage
R@10
MRR@10
nDCG@10
ObliQA (ADGM rulebooks)
LegalBERT (baseline)
0.256
0.158
0.181
+ entailment tuning (ET)
0.495
0.322
0.363
+ contrastive tuning (CT)
0.732
0.518
0.569
+ BM25 fusion (hybrid)
0.774
0.594
0.638
Open Australian LegalQA
Table 3: Staged retriever ablation. Each stage is initialized from the previous one; the hybrid adds BM25 at α=0.8 .
Retriever
Type
R@10
MRR@10
nDCG@10
BM25
lexical
0.678
0.504
0.546
BGE-M3
dense
0.711
0.544
0.584
BGE-base-en-v1.5
dense
0.719
0.549
0.590
E5-base-v2
dense
0.721
0.557
0.597
E5-large-v2
dense
0.758
0.595
0.635
Hybrid ET+CT (ours)
lexical+dense
0.774
0.594
0.638
Table 4: ObliQA retrieval against standard baselines, full validation split (3,733 question–passage pairs, 13,705-passage corpus).
Model
Strategy
Es↑
Cs↓
OCs↑
RePASs ↑
Qwen2.5-7B
zero-shot
0.894
0.189
0.301
0.668
few-shot
0.971
0.167
0.371
0.725
few-shot + CoT
0.957
0.154
0.359
0.721
RAFT-LoRA
0.926
0.157
0.409
0.725
Teuken-7B
zero-shot
0.944
0.256
0.332
0.673
v0.6
few-shot
0.928
0.200
0.259
0.662
Table 5: Answer generation on ObliQA, 150 evaluation questions, hybrid retriever context. Best RePASs per model in bold.
Model
Strategy
Es↑
Cs↓
OCs↑
RePASs ↑
Qwen2.5-7B
zero-shot
0.964
0.414
0.377
0.642
few-shot
0.964
0.402
0.350
0.637
few-shot + CoT
0.952
0.492
0.242
0.567
RAFT-LoRA †
0.957
0.451
0.238
0.581
Teuken-7B
zero-shot
0.943
0.552
0.288
0.559
v0.6
few-shot
0.941
0.578
0.212
0.525
Table 6: Answer generation on Open Australian LegalQA, 150 evaluation questions. † marks adapters trained on ObliQA and applied without retraining.
Model
Setting
Es↑
Cs↓
OCs↑
RePASs ↑
Qwen2.5-7B
closed-book
0.883
0.081
0.169
0.657
RAG (zero-shot)
0.894
0.189
0.301
0.668
Teuken-7B v0.6
closed-book
0.836
0.096
0.247
0.662
RAG (zero-shot)
0.944
0.256
0.332
0.673
Table 7: Closed-book control on ObliQA. Scores are computed against retrieved passages that the closed-book models did not receive. Closed-book runs cover 132 (Qwen) and 89 (Teuken) validation questions.