Posterior Regimes and Latent Deception: Variational Bayesian Inference in Hidden Markov Models for Sequential Fraud Detection in Financial Transactions
Authors: Joseph Uririoghene Obukofe, Anthony O'Hare, Chioma Sandra Dike
Organizations: Computing Science and Mathematics, Faculty of Natural Sciences, University of Stirling, Stirling FK9 4LA, Scotland, United Kingdom · University of Nottingham, Nottingham, NG7 2RD, United Kingdom
We present a three-tier progression of Hidden Markov Models: maximum-likelihood (Baum-Welch), variational Bayesian (VBEM), and a neural variational extension (Neural VBEM), that model each customer's transaction history as a trajectory through a small number of latent behavioural regimes, one of which is empirically identified as fraud-associated. The Neural VBEM HMM replaces the fixed Gaussian-multinomial emission family with a learned encoder, compressing a 741-dimensional transaction representation into a 64-dimensional latent space in which the VBEM HMM's posterior operates; a UMAP projection of this space reveals that the discovered regimes are not discrete clusters but ordered segments of a single continuous behavioural manifold, with confirmed fraud concentrated at its extreme. We show that the model's natural output, that is, the posterior probability of regime membership, is routinely mistaken for a fraud probability, and quantify the resulting miscalibration (the regime-membership interpretation error, MRIE); a corrected posterior-predictive score, closes most of this gap. We further distinguish batch (smoothed) inference, which uses look-ahead unavailable at deployment time, from filtered (forward-only) inference, and report both. On IEEE-CIS transaction data, the neural tier achieves a 14.4× fraud enrichment in its identified regime; while its AUPRC trails a discriminative XGBoost baseline, we show this gap is structural and not incidental, and argue the model is best positioned as a calibrated triage and interpretability layer rather than a drop-in ranking replacement.
Figures & tables
Model
K∗
Occupancy
Fraud rate
Enrichment
Baum-Welch HMM
8
1.77%
19.4%
5.48×
VBEM HMM
7
3.78%
25.3%
7.16×
Neural VBEM HMM ( dz=64 )
9
3.8%
50.9%
14.43×
Table 1: Selected model order and fraud-state diagnostics for each HMM tier under batch posteriors for engineered features.
Model
Score
AUPRC
KS
ECE
Baum-Welch HMM
raw / batch
0.0775
0.2667
0.0414
corrected / batch
0.0853
0.3883
≈0
corrected / filtered
0.0853
0.3883
≈0
VBEM HMM
raw / batch
0.0667
0.2151
0.0685
corrected / batch
0.0806
0.2841
≈0
corrected / filtered
0.0806
0.2841
≈0
Table 2: HMM tier predictive performance under three scoring conditions. Raw/batch is γt(k∗) , smoothed posteriors; corrected/batch and corrected/filtered use Fcorrected=∑kγt(k)ηk ( Section 3.6 ) under smoothed and forward-only posteriors respectively.
Model
AUPRC
KS
ECE
Isolation Forest
0.098
0.365
0.458
Neural VBEM HMM
0.199
0.471
0.049
XGBoost
0.514
0.563
0.019
Table 3: Predictive performance against non-sequential baselines on held-out evaluation set. Neural VBEM figure is the raw/batch score, γt(k∗) ; see note below.
Figure 1 : UMAP projection of the Neural VBEM HMM ( dz=64 , K∗=9 ) latent embeddings. Panel (i): each transaction coloured by its hard state assignment, fraud-associated state k∗=4 highlighted in red. Panel (ii): the same projection coloured by ground-truth fraud label. Confirmed fraud concentrates at the same terminal tip, density diminishing continuously toward the legitimate-dominated body.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Form
Description
D
Set
Full dataset of N labelled transactions.
N
Scalar integer
Total number of transactions.
U
Scalar integer
Total number of unique customers.
u
Index
Customer index, u=1,…,U .
n
Index
Global transaction index, n=1,…,N .
t
Index
Chronological index within customer u ’s sequence, t=1,…,Tu .
Appendix
Table 4 : Dataset and Sequence Notation
Symbol
Form
Description
K , K∗
Scalar integers
Candidate and selected model order.
k,i,j
State indices
General, from-, and to-state indices over 1,…,K .
qu,t , qu
Latent, Latent path
Hidden state at t ; full hidden state sequence.
k∗
Scalar integer
Index of the fraud-associated state, identified post-training as argmaxkηk .
θ
Parameter set
All HMM parameters: π,A,μk,Σk,πk(j) .
π , πk
Vector, scalar
Initial state distribution; probability of starting in state k .
Appendix
Table 5 : Hidden Markov Model Notation
Symbol
Form
Description
αt(k) , α
Scalar, matrix
Forward variable; full forward matrix for a sequence.
βt(k) , β
Scalar, matrix
Backward variable; full backward matrix for a sequence.
γt(k) , γ
Scalar, matrix
Batch (smoothed) state-occupancy posterior; full matrix for a sequence.
γtfilt(k)
Scalar
Filtered (forward-only) state-occupancy posterior: conditions on xu,1:t only, no backward pass.
Working entirely on topologically anonymized embeddings, we perform fraud detection using iterative rounds of unsupervised filtering followed by supervised sniping. The result is an ultra-low latency privacy--preserving triage that allows institutions to flag suspicious activity without compromising Personally Identifiable Information.
Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data. We present an end-to-end pipeline for customer-level mule detection comprising three stages: (1) a LightGBM classifier trained on 280 engineered features spanning transaction patterns, account demographics, network topology, and temporal behaviour; (2) a TreeSHAP attribution layer that decomposes each prediction into feature contributions; and (3) a large language model (LLM) module that converts SHAP attributions into analyst-facing natural-language narratives. We evaluate across three open-weight LLM families and assess explanation quality through analyst feedback. In a live production deployment, the system achieves a yield rate of 89%, up from 61% under the incumbent rule-based system, with monthly alert volume expanding from 211 to 302, reflecting broader true-positive coverage rather than increased noise. This corresponds to a 60% incremental adverse detection beyond existing review workflows, substantially outperforming the rule-based approach. Qualitative feedback from analysts indicates that LLM-generated narratives reduce cognitive load during alert triage. We further discuss implications of deploying LLM-augmented explainability in regulated financial environments.
Real-time payment fraud detection is a non-stationary streaming prediction problem: adversaries adapt before supervised labels mature, and localized burst attacks can cause losses before retraining. Production systems typically rely on tabular classifiers and rules, which can struggle to capture these emerging sequential patterns before periodic retraining occurs. We present SR-Fraud, an outcome-supervised reflective LLM framework that decouples request-time decisions from offline adaptation. A frozen, stateless agent scores each transaction from a Hybrid Episodic Window to track behavioral shifts, while an offline reflection agent proposes boundary hypotheses from matured errors. A deterministic verifier then admits only supported hypotheses into an executable knowledge state. On a production payment-fraud benchmark, SR-Fraud improves all detection metrics over its frozen decision agent, obtains higher point estimates than static and periodically retrained CatBoost, and detects an emerging fraud burst.