Posterior Regimes and Latent Deception: Variational Bayesian Inference in Hidden Markov Models for Sequential Fraud Detection in Financial Transactions
Organizations: Computing Science and Mathematics, Faculty of Natural Sciences, University of Stirling, Stirling FK9 4LA, Scotland, United Kingdom · University of Nottingham, Nottingham, NG7 2RD, United Kingdom
Abstract
We present a three-tier progression of Hidden Markov Models: maximum-likelihood (Baum-Welch), variational Bayesian (VBEM), and a neural variational extension (Neural VBEM), that model each customer's transaction history as a trajectory through a small number of latent behavioural regimes, one of which is empirically identified as fraud-associated. The Neural VBEM HMM replaces the fixed Gaussian-multinomial emission family with a learned encoder, compressing a 741-dimensional transaction representation into a 64-dimensional latent space in which the VBEM HMM's posterior operates; a UMAP projection of this space reveals that the discovered regimes are not discrete clusters but ordered segments of a single continuous behavioural manifold, with confirmed fraud concentrated at its extreme. We show that the model's natural output, that is, the posterior probability of regime membership, is routinely mistaken for a fraud probability, and quantify the resulting miscalibration (the regime-membership interpretation error, MRIE); a corrected posterior-predictive score, closes most of this gap. We further distinguish batch (smoothed) inference, which uses look-ahead unavailable at deployment time, from filtered (forward-only) inference, and report both. On IEEE-CIS transaction data, the neural tier achieves a 14.4 fraud enrichment in its identified regime; while its AUPRC trails a discriminative XGBoost baseline, we show this gap is structural and not incidental, and argue the model is best positioned as a calibrated triage and interpretability layer rather than a drop-in ranking replacement.
Figures & tables
| Model | Occupancy | Fraud rate | Enrichment | |
|---|---|---|---|---|
| Baum-Welch HMM | 8 | 1.77% | 19.4% | |
| VBEM HMM | 7 | 3.78% | 25.3% | |
| Neural VBEM HMM ( ) | 9 | 3.8% | 50.9% |
| Model | Score | AUPRC | KS | ECE |
|---|---|---|---|---|
| Baum-Welch HMM | raw / batch | 0.0775 | 0.2667 | 0.0414 |
| corrected / batch | 0.0853 | 0.3883 | ||
| corrected / filtered | 0.0853 | 0.3883 | ||
| VBEM HMM | raw / batch | 0.0667 | 0.2151 | 0.0685 |
| corrected / batch | 0.0806 | 0.2841 | ||
| corrected / filtered | 0.0806 | 0.2841 |
| Model | AUPRC | KS | ECE |
|---|---|---|---|
| Isolation Forest | 0.098 | 0.365 | 0.458 |
| Neural VBEM HMM | 0.199 | 0.471 | 0.049 |
| XGBoost | 0.514 | 0.563 | 0.019 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Form | Description |
|---|---|---|
| Set | Full dataset of labelled transactions. | |
| Scalar integer | Total number of transactions. | |
| Scalar integer | Total number of unique customers. | |
| Index | Customer index, . | |
| Index | Global transaction index, . | |
| Index | Chronological index within customer ’s sequence, . |
| Symbol | Form | Description |
|---|---|---|
| , | Scalar integers | Candidate and selected model order. |
| State indices | General, from-, and to-state indices over . | |
| , | Latent, Latent path | Hidden state at ; full hidden state sequence. |
| Scalar integer | Index of the fraud-associated state, identified post-training as . | |
| Parameter set | All HMM parameters: . | |
| , | Vector, scalar | Initial state distribution; probability of starting in state . |
| Symbol | Form | Description |
|---|---|---|
| , | Scalar, matrix | Forward variable; full forward matrix for a sequence. |
| , | Scalar, matrix | Backward variable; full backward matrix for a sequence. |
| , | Scalar, matrix | Batch (smoothed) state-occupancy posterior; full matrix for a sequence. |
| Scalar | Filtered (forward-only) state-occupancy posterior: conditions on only, no backward pass. | |
| , | Scalar, tensor | Transition posterior; full tensor for a sequence. |
| Operator | Element-wise multiplication. |
| Symbol | Form | Description |
|---|---|---|
| , | Scalars | Marginal and joint sequence probabilities. |
| , , | Distributions, scalar | Prior, true posterior, and model evidence. |
| , | Distributions | Variational posterior; its mean-field factors. |
| Functional | Kullback-Leibler divergence. | |
| Operator | Expectation under . | |
| Scalar | Evidence Lower Bound. |
| Symbol | Form | Description |
|---|---|---|
| Function | Neural encoder, . | |
| , , | Vector, integers | Latent embedding; latent and hidden dimensions ( , ). |
| , | Integer, matrix | Embedding dimension and table for categorical feature . |
| Vector | Normalised continuous features fed to the encoder MLP (LayerNorm output). | |
| , | Operator, activation | Layer Normalisation; Gaussian Error Linear Unit. |
| , , | Vectors | Continuous-pathway, categorical-pathway, and fused hidden representations. |
| Symbol | Form | Description |
|---|---|---|
| Vector, matrix | Population-level NIW prior (moment-matched). | |
| Vector, matrix | Customer- converged posterior statistics feeding the population prior. | |
| Scalar | Cold-start fraud score for new customer . |
| Symbol | Form | Description |
|---|---|---|
| , | Sets | Candidate model orders ; the eligible subset. |
| , | Scalars | Effective count and fractional occupancy of state under order . |
| , | Scalars | Fraud-weighted effective count; empirical fraud rate of state under order . |
| , | Integer, scalar | Fraud-associated state under order ; its enrichment ratio. |
| , | Scalars | Population fraud rate ( ); parsimony tolerance ( ). |
| Notation | Shorthand for “fraud”. |
| Symbol | Form | Description |
|---|---|---|
| Vector, | Transaction consequence: the real-time belief state, updated one transaction at a time from ; coincides with . | |
| Scalar | Batch fraud score, . | |
| Scalar | Real-time fraud score, . | |
| Scalar | Posterior-predictive corrected fraud score, . | |
| Scalar | Mean Regime Interpretation Error, . |
| K | Occup. (p) | Occup. (e) | F. Rate (p) | F. Rate (e) | Enrich. (p) | Enrich. (e) |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| K | Occup. (p) | Occup. (e) | F. Rate (p) | F. Rate (e) | Enrich. (p) | Enrich. (e) |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| K | Occup. (p) | Occup. (e) | F. Rate (p) | F. Rate (e) | Enrich. (p) | Enrich. (e) |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| K | Occup. (p) | Occup. (e) | F. Rate (p) | F. Rate (e) | Enrich. (p) | Enrich. (e) |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| K | Occup. (p) | Occup. (e) | F. Rate (p) | F. Rate (e) | Enrich. (p) | Enrich. (e) |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| K | Occup. (p) | Occup. (e) | F. Rate (p) | F. Rate (e) | Enrich. (p) | Enrich. (e) |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| K | Occup. (p) | Occup. (e) | F. Rate (p) | F. Rate (e) | Enrich. (p) | Enrich. (e) |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| Model | Occupancy | Fraud rate | Enrichment | |
|---|---|---|---|---|
| Baum-Welch HMM | 8 | |||
| VBEM HMM | - | - | - | - |
| Neural VBEM ( ) | 8 | |||
| Neural VBEM ( ) | 10 | |||
| Neural VBEM ( ) | 10 |
| Model | Occupancy | Fraud rate | Enrichment | |
|---|---|---|---|---|
| Baum-Welch HMM | 8 | |||
| VBEM HMM | 7 | |||
| Neural VBEM ( ) | 8 | |||
| Neural VBEM ( ) | 9 | |||
| Neural VBEM ( ) | 9 |
| Component | Value |
|---|---|
| Encoder architecture | |
| Hidden dimension | 512 (fixed) |
| Latent dimension | (swept) |
| Categorical embedding dim. | |
| Continuous pathway | LayerNorm 2-layer MLP, GELU |
| Dropout | 0.1 |