Machine learning for journal entry testing: A type-aware evaluation of anomaly detectors under a review budget
Organizations: Saarland University, Campus, 66123, Saarbrücken, Germany · German Research Center for Artificial Intelligence (DFKI), Institute for Information Systems (IWi), Campus D3.2, 66123, Saarbrücken, Germany · Leipzig University, Grimmaische Straße 12, 04109, Leipzig, Germany
Abstract
Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, type coverage and first-hit rank. We evaluate nine unsupervised detectors, a supervised reference and feedback-driven Deep Semi-Supervised Anomaly Detection (DeepSAD) on four real client ledgers with injected typed anomalies and a public synthetic ledger. On the largest client ledger, principal component analysis (PCA), an autoencoder (AE) and a variational autoencoder (VAE) each place on average 98 anomalies among the first 100 postings, but at least 95.8 belong to one type. FSR instead favours a nearest-neighbour (kNN) detector and changes the top-ranked detector on three of four client ledgers. Representation also matters: one-hot encoding exposes unseen accounts, whereas frequency encoding leaves unseen contra accounts largely undetected. On the public ledger, the Histogram-Based Outlier Score (HBOS) and Empirical Cumulative Distribution-Based Outlier Detection (ECOD) reach all eight markings within 1,386 entries, whereas kNN, the hit leader at 1,000 entries, first reaches cross-linked clearing at rank 4,641, and the supervised row-level reference misses this marking within 1,000 entries. There, the adaptive DeepSAD review protocol raises mean hits per 100 reviews from 40.0 to 68.3 but type coverage only from 2.7 to 3.0. These findings show that high hit rates can conceal systematic blind spots and suggest that feedback can reinforce existing detection patterns without broadening anomaly coverage.
Figures & tables
| Measure | Bound to the budget | Credits breadth over types | Adjusts for type size | Caps over-representation | Order within the budget | Risk weights |
| yes | no | no | no | no | no | |
| Intent-aware precision | yes | no (orders as ) | no | no | no | yes |
| Macro recall at | yes | yes | inversely (favours small types) | only at | no | possible |
| yes | yes (binary) | not needed | at one hit | no | possible | |
| -nDCG at | yes | yes | no | soft decay | yes | in intent-aware variant |
| yes | yes | yes (fair shares) | at the fair share | no (left to ) | yes |
| Client ledgers | Public | ||||
| Property | C | A | B | D | ledger |
| Scored unit | posting | posting | posting | posting | line |
| Evaluated unit | posting | posting | posting | posting | entry |
| Scored rows | 3,461 | 5,080 | 17,505 | 279,975 | 256,232 |
| Evaluated units | 3,461 | 5,080 | 17,505 | 279,975 | 102,429 |
| Fiscal years | 1 | 1 | 1 | 1 | 2 |
| (a) Types injected into the client ledgers | ||
| Type (Foorthuis subtype) | Modification of a sampled posting | : C / A / B / D |
| T1 extreme amount (Ia) | amount z-score set to with a random sign; | 22 / 31 / 105 / 1,680 |
| T2 extreme payment period (Ia) | payment period set to , rows with an observed period only; | 10 / 15 / 53 / 840 |
| T3 unseen contra account (IIa) | contra account replaced by a fresh id, one distinct id per row | 4 / 6 / 21 / 336 |
| T4 unseen account (IIa) | account replaced by a fresh id, one distinct id per row | 4 / 6 / 21 / 336 |
| T5 unusual account pair (Va) | (account, contra account) replaced by a pair that never occurs in the ledger and whose members each occur at least times; | 12 / 18 / 63 / 1,008 |
| Detector | Family | Input | Key settings |
| IF | isolation, tree ensemble | 200 trees, 256 samples per tree | |
| LOF | local density | 20 neighbours, the row itself excluded | |
| kNN | distance to neighbours | distance to the fifth nearest other row | |
| HBOS | per-field histogram | 10 bins, , tolerance 0.5 | |
| ECOD | empirical tail probability | no free parameters; computed column-chunked, scores identical to the library | |
| PCA | linear projection | components explaining 90 percent of the variance; score = sum of the distances to the retained eigenvectors divided by their explained-variance ratios |
| Client ledgers (five seeds) | Public | ||||
| Detector | C | A | B | D | ledger |
| (a) P@100 | |||||
| IF | 0.04 (0.01) | 0.03 (0.02) | 0.04 (0.04) | 0.11 (0.07) | 0.08 (0.04) |
| LOF | 0.06 (0.03) | 0.01 (0.01) | 0.00 (0.00) | 0.11 (0.03) | 0.11 |
| kNN | 0.40 (0.02) | 0.29 (0.04) | 0.42 (0.03) | 0.34 (0.04) | 0.17 |
| HBOS | 0.03 (0.01) | 0.02 (0.02) | 0.03 (0.03) | 0.04 (0.01) | 0.04 |
| Best by hits | Best by FSR@ | ||||||
| Ledger ( ) | Detector | P@ | TC@ | FSR@ | Detector | FSR@ | |
| C (100) | kNN | 0.40 | n/d | 0.57 | VAE | 0.65 | 0.83 |
| A (100) | kNN | 0.29 | n/d | 0.29 | AE | 0.38 | 0.65 |
| B (100) | kNN | 0.42 | n/d | 0.31–0.32 | kNN | 0.31–0.32 | 0.97–0.99 |
| D (100) | PCA, AE, VAE | 0.98 | n/d | 0.200–0.226 | kNN | 0.284–0.344 | 0.598–0.625 |
| Public (100) | AE-emb | 0.40 | 2.7 | 0.22 | AE-emb | 0.22 | 1.00 |
| Leader by | Kendall’s | |||||
| Ledger ( ) | Cap | P@ | Macro | FSR@ | P, macro | macro, FSR |
| C (100) | kNN | VAE | VAE | 0.83 | 1.00 | |
| A (100) | kNN | AE | AE | 0.65 | 1.00 | |
| B (100) | 20 | kNN | VAE, AE | kNN | 0.74 | 0.69 |
| D (100) | 20 | PCA, AE, VAE | PCA, AE, VAE | kNN | 0.81 | 0.42 |
| Public (100) | 12.5 | AE-emb | AE-emb | AE-emb | 0.94 | 0.94 |
| (a) Anomalies found in 100 reviews | ||||||||
| DeepSAD, after reviews | AE-emb, first 100 | Most hits, first 100 | ||||||
| Ledger | 25 | 50 | 75 | 100 | Found | Gain | Detector | Found |
| C | 3.0 | 21.0 | 26.2 | 29.4 (4.6) | 12 (2) | 17.4 | kNN | 40 |
| A | 8.8 | 29.0 | 40.6 | 47.0 (3.0) | 19 (1) | 28.0 | kNN | 29 |
| B | 8.0 | 23.4 | 39.4 | 57.6 (11.7) | 23 (7) | 34.6 | kNN | 42 |
| D | 10.2 | 29.4 | 52.8 | 76.0 (8.9) | 27 (11) | 49.0 | PCA, AE, VAE | 98 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
| , | number of units; label of unit (1 anomalous, 0 normal) |
| , | set of anomaly types (T1 to T5, or markings M1 to M8) and its size |
| , | anomalies of type and their number |
| , | score of unit ; unit at rank (descending score, ties by position) |
| , | review budget; budget as a share of the ledger, |
| , | the first units of the ranking; number of anomalies of type among them |
| Seed | List | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | F@25 | F@100 | TC@100 | FSR@100 |
| 0 | AE-emb, first 100 | 0 | 0 | 0 | 0 | 1 | 0 | 18 | 0 | 6 | 19 | 2 | 0.135 |
| 0 | DeepSAD, reviewed | 5 | 0 | 0 | 0 | 6 | 0 | 60 | 0 | 6 | 71 | 3 | 0.235 |
| 0 | Control, reviewed | 25 | 0 | 0 | 0 | 1 | 0 | 35 | 0 | 6 | 61 | 3 | 0.260 |
| 1 | AE-emb, first 100 | 26 | 0 | 0 | 0 | 6 | 0 | 5 | 0 | 15 | 37 | 3 | 0.235 |
| 1 | DeepSAD, reviewed | 58 | 0 | 0 | 0 | 1 | 0 | 5 | 0 | 15 | 64 | 3 | 0.185 |
| 1 | Control, reviewed | 28 | 0 | 0 | 0 | 2 | 0 | 5 | 0 | 15 | 35 | 3 | 0.195 |
| (a) Point estimate [95 percent interval] | |||||
| Detector | P@100 | FSR@100 | P@1,000 | FSR@1,000 | TC@1,000 |
| IF | 0.08 [0.03, 0.14] | 0.08 [0.03, 0.14] | 0.083 [0.067, 0.100] | 0.111 [0.087, 0.136] | 7 [6, 7] |
| LOF | 0.11 [0.05, 0.18] | 0.11 [0.05, 0.13] | 0.074 [0.059, 0.092] | 0.128 [0.112, 0.145] | 4 [3, 4] |
| kNN | 0.17 [0.10, 0.26] | 0.17 [0.10, 0.26] | 0.215 [0.189, 0.240] | 0.301 [0.275, 0.323] | 6 [5, 7] |
| HBOS | 0.04 [0.01, 0.08] | 0.04 [0.01, 0.08] | 0.067 [0.052, 0.083] | 0.104 [0.077, 0.131] | 7 [5, 7] |
| ECOD | 0.04 [0.01, 0.08] | 0.04 [0.01, 0.08] | 0.083 [0.065, 0.100] | 0.136 [0.110, 0.158] | 6 [5, 6] |
| (a) Client ledgers, | |||||
| Ledger | P@100 leader | FSR leader, equal weights | Reversal fixed (of 11 schemes) | Exceptions | Random vectors with reversal fixed |
| C | kNN | VAE | 10 | T4 halved | 91% |
| A | kNN | AE | 11 | none | 89% |
| B | kNN | kNN (no reversal) | 2 | reversal to VAE when T1 is halved or T4 doubled | 26% |
| D | PCA, AE, VAE | kNN | 9 | T1 halved, T4 doubled (undetermined) | 63% |
| Detector | P@25 | P@50 | P@100 | AP | ROC-AUC |
| (a) Ledger C: 3,461 postings, 52 injected anomalies, 5 seeds | |||||
| IF | 0.04 (0.03) | 0.04 (0.02) | 0.04 (0.01) | 0.04 (0.01) | 0.67 (0.03) |
| LOF | 0.18 (0.07) | 0.11 (0.05) | 0.06 (0.03) | 0.11 (0.02) | 0.83 (0.04) |
| kNN | 0.65 (0.02) | 0.64 (0.02) | 0.40 (0.02) | 0.52 (0.02) | 0.98 (0.00) |
| HBOS | 0.02 (0.04) | 0.02 (0.03) | 0.03 (0.01) | 0.03 (0.01) | 0.66 (0.03) |
| ECOD | 0.02 (0.04) | 0.08 (0.02) | 0.08 (0.01) | 0.05 (0.01) | 0.79 (0.02) |
| Area | Deviation or implementation note |
| Injection | Ledger A, easy level, %: the pool holds 19 valid account pairs for the 36 T5 rows the cell needs, so 17 rows reuse a pair. The protocol permits drawing with replacement once the pool is exhausted; the reuse is recorded per cell. |
| Typology | The protocol called T5 a multidimensional mixed-data anomaly (Foorthuis Type VI). T5 changes two categorical fields only, so it is classified as Type V, subtype Va. |
| PCA | The score is PyOD’s own: the sum of the distances between the standardised row and the retained eigenvectors, each divided by its explained-variance ratio, not a literal reconstruction error. The covariance eigensolver replaces the full singular value decomposition (same components, scores agree to relative or better, identical P@ and AP). |
| ECOD | PyOD’s algorithm recomputed column chunk by column chunk, because the library implementation exhausts memory on ledger D; maximum difference to the library on ledgers C, A and B. Used on every ledger. |
| Networks on ledger D | AE, VAE and AE-emb train with batch 256 for 30 epochs instead of batch 64 for 100; the protocol grants this for AE-emb, and a runtime probe extended it to the two PyOD networks. |
| Execution on ledger D | One thread per process for the linear algebra libraries and PyTorch; every detector runs in a checkpointing worker process; LOF and kNN are fitted on all 279,975 rows, the subsampling rule being implemented but never triggered. Scores are unaffected. |
| Role | Client ledgers | Public ledger |
| Scored row | one posting | one posting line |
| Evaluated unit | one posting | one journal entry; its score is the maximum over its lines, and labels are constant within an entry |
| Account | Kontonummer | gl_account |
| Contra account | Gegenkonto | derived: the account of the largest-amount line on the opposite side of the same entry, 21 levels |
| Amount | Umsatz_mod_z , the modified z-score supplied with the data | amount ; and standardised in the encodings |
| Payment period | Zahlungszeitraum | entry lag in days, entered_date minus document_date ; used in the encodings |
| Detector | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | All |
| IF | 1 | 51 | 626 | 309 | 46 | 217 | 3,415 | 774 | 3,415 |
| min | 1 | 37 | 327 | 74 | 12 | 140 | 2,157 | 661 | 2,157 |
| max | 1 | 62 | 1,108 | 654 | 92 | 279 | 4,486 | 944 | 4,486 |
| LOF | 239 | 3,116 | 472 | 11,107 | 192 | 3,098 | 1 | 6,895 | 11,107 |
| kNN | 43 | 185 | 824 | 15 | 70 | 1,001 | 53 | 4,641 | 4,641 |
| HBOS | 1 | 429 | 1,386 | 661 | 456 | 479 | 30 | 703 | 1,386 |
| Detector | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 |
| 120 | 60 | 120 | 120 | 100 | 250 | 60 | 20 | |
| (a) Recall within the first 100 entries | ||||||||
| IF | 0.02 | 0.04 | 0.00 | 0.00 | 0.02 | 0.00 | 0.00 | 0.00 |
| LOF | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.18 | 0.00 |
| kNN | 0.02 | 0.00 | 0.00 | 0.07 | 0.03 | 0.00 | 0.07 | 0.00 |
| HBOS | 0.02 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.03 | 0.00 |
| Detector | T1 | T2 | T3 | T4 | T5 | FSR@100 |
| (a) Ledger C | ||||||
| 22 | 10 | 4 | 4 | 12 | ||
| IF | 0.12 | 0.16 | 0.05 | 0.00 | 0.00 | 0.07 |
| LOF | 0.22 | 0.00 | 0.00 | 0.30 | 0.03 | 0.11 |
| kNN | 1.00 | 1.00 | 0.15 | 0.10 | 0.58 | 0.57 |
| HBOS | 0.06 | 0.04 | 0.05 | 0.10 | 0.03 | 0.06 |
| Detector | T1 | T2 | T3 | T4 | T5 | FSR@100 |
| (c) Ledger B | ||||||
| 105 | 53 | 21 | 21 | 63 | ||
| IF | 0.02 | 0.04 | 0.00 | 0.00 | 0.00 | 0.04 |
| LOF | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| kNN | 0.29 | 0.20 | 0.01 | 0.00 | 0.01 | 0.31 to 0.32 |
| HBOS | 0.02 | 0.01 | 0.00 | 0.02 | 0.01 | 0.03 |
| (a) Precision at review budgets | ||||
| Detector | P@25 | P@50 | P@100 | P@1,000 |
| IF | 0.09 (0.02) | 0.07 (0.03) | 0.08 (0.04) | 0.067 (0.024) |
| LOF | 0.12 | 0.14 | 0.11 | 0.074 |
| kNN | 0.08 | 0.12 | 0.17 | 0.215 |
| HBOS | 0.08 | 0.08 | 0.04 | 0.067 |
| ECOD | 0.08 | 0.08 | 0.04 | 0.083 |
| J1 | J2 | J3 | J4 | |
| Rule | top-n amounts | high cash | delayed entry | weekend or non-working hours |
| Entries | 2,000 | 402 | 8,207 | 16,493 |
| Anomalies | 196 | 190 | 310 | 415 |
| Rule precision | 0.098 | 0.473 | 0.038 | 0.025 |
| (a) P@100 inside the population | ||||
| IF | 0.40 | 0.53 | 0.06 | 0.08 |
| Client ledgers | Public | ||||
| Detector | C | A | B | D | ledger |
| IF | 14.1 | 14.7 | 0.6 | 9.8 | 5.1 |
| LOF | 0.1 | 0.1 | 1.5 | 1,022.6 | 478.2 |
| kNN | 0.0 | 0.1 | 1.5 | 1,029.0 | 467.1 |
| HBOS | 1.3 | 1.3 | 0.2 | 10.5 | 4.6 |
| ECOD | 0.1 | 0.1 | 0.4 | 25.1 | 4.8 |
| % | % | % | |||||||
| Detector | easy | medium | hard | easy | medium | hard | easy | medium | hard |
| (a) Ledger C | |||||||||
| IF | 0.02 | 0.01 | 0.01 | 0.04 | 0.04 | 0.04 | 0.09 | 0.09 | 0.08 |
| LOF | 0.04 | 0.03 | 0.02 | 0.11 | 0.06 | 0.03 | 0.07 | 0.04 | 0.02 |
| kNN | 0.15 | 0.14 | 0.14 | 0.43 | 0.40 | 0.37 | 0.75 | 0.69 | 0.64 |
| HBOS | 0.01 | 0.01 | 0.01 | 0.03 | 0.03 | 0.03 | 0.05 | 0.06 | 0.06 |