Machine learning for journal entry testing: A type-aware evaluation of anomaly detectors under a review budget
Authors: Jan Gronewald, Michel Scherer, Nijat Mehdiyev
Organizations: Saarland University, Campus, 66123, Saarbrücken, Germany · German Research Center for Artificial Intelligence (DFKI), Institute for Information Systems (IWi), Campus D3.2, 66123, Saarbrücken, Germany · Leipzig University, Grimmaische Straße 12, 04109, Leipzig, Germany
Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, type coverage and first-hit rank. We evaluate nine unsupervised detectors, a supervised reference and feedback-driven Deep Semi-Supervised Anomaly Detection (DeepSAD) on four real client ledgers with injected typed anomalies and a public synthetic ledger. On the largest client ledger, principal component analysis (PCA), an autoencoder (AE) and a variational autoencoder (VAE) each place on average 98 anomalies among the first 100 postings, but at least 95.8 belong to one type. FSR instead favours a nearest-neighbour (kNN) detector and changes the top-ranked detector on three of four client ledgers. Representation also matters: one-hot encoding exposes unseen accounts, whereas frequency encoding leaves unseen contra accounts largely undetected. On the public ledger, the Histogram-Based Outlier Score (HBOS) and Empirical Cumulative Distribution-Based Outlier Detection (ECOD) reach all eight markings within 1,386 entries, whereas kNN, the hit leader at 1,000 entries, first reaches cross-linked clearing at rank 4,641, and the supervised row-level reference misses this marking within 1,000 entries. There, the adaptive DeepSAD review protocol raises mean hits per 100 reviews from 40.0 to 68.3 but type coverage only from 2.7 to 3.0. These findings show that high hit rates can conceal systematic blind spots and suggest that feedback can reinforce existing detection patterns without broadening anomaly coverage.
Figures & tables
Figure 1: Type-aware budget evaluation. Four client ledgers with injected types T1 to T5 and a public ledger with markings M1 to M8 are ranked by nine unsupervised detectors and a supervised LightGBM reference; an auditor reviews the first k units. The conventional view counts the anomalies among them, the type-aware view asks which types they are. The example lists are schematic.
Measure
Bound to the budget
Credits breadth over types
Adjusts for type size
Caps over-representation
Order within the budget
Risk weights
P@k
yes
no
no
no
no
no
Intent-aware precision
yes
no (orders as P@k )
no
no
no
yes
Macro recall at k
yes
yes
inversely (favours small types)
only at nt
no
possible
TC@k
yes
yes (binary)
not needed
at one hit
no
possible
α -nDCG at k
yes
yes
no
soft decay
yes
in intent-aware variant
FSR@k
yes
yes
yes (fair shares)
at the fair share
no (left to FHRt )
yes
Table 1: Properties of hit-based and type-aware measures for a review budget of k units with anomaly types of unequal size. Intent-aware precision is precision under the intent-aware construction of Agrawal et al. (2009) , with anomaly types as intents; macro recall is the mean of Rt@k over the types.
Client ledgers
Public
Property
C
A
B
D
ledger
Scored unit
posting
posting
posting
posting
line
Evaluated unit
posting
posting
posting
posting
entry
Scored rows
3,461
5,080
17,505
279,975
256,232
Evaluated units
3,461
5,080
17,505
279,975
102,429
Fiscal years
1
1
1
1
2
Table 2: Profile of the four client ledgers and the public synthetic ledger. Counts of accounts, contra accounts, tax codes and users are numbers of distinct values. A duplicate row repeats all feature values of another row. Anomalous units are the injected anomalies of the core grid (client ledgers) and the generator-labelled journal entries (public ledger).
(a) Types injected into the client ledgers
Type (Foorthuis subtype)
Modification of a sampled posting
nt : C / A / B / D
T1 extreme amount (Ia)
amount z-score set to ±fP99.9(∣z∣) with a random sign; f=5,2,1
22 / 31 / 105 / 1,680
T2 extreme payment period (Ia)
payment period set to fP99.9 , rows with an observed period only; f=5,2,1
10 / 15 / 53 / 840
T3 unseen contra account (IIa)
contra account replaced by a fresh id, one distinct id per row
4 / 6 / 21 / 336
T4 unseen account (IIa)
account replaced by a fresh id, one distinct id per row
4 / 6 / 21 / 336
T5 unusual account pair (Va)
(account, contra account) replaced by a pair that never occurs in the ledger and whose members each occur at least F times; F=50,20,5
12 / 18 / 63 / 1,008
Table 3: Anomaly types and markings. (a) Types injected into the client ledgers, with their subtype in the typology of Foorthuis (2021) , the difficulty parameter for the levels easy, medium and hard, and the number nt of injected rows in the core grid. (b) Markings of the public ledger, the rule population into which each was injected, and nt .
Detector
Family
Input
Key settings
IF
isolation, tree ensemble
Ebase
200 trees, 256 samples per tree
LOF
local density
Ebase
20 neighbours, the row itself excluded
kNN
distance to neighbours
Ebase
distance to the fifth nearest other row
HBOS
per-field histogram
Ebase
10 bins, α=0.1 , tolerance 0.5
ECOD
empirical tail probability
Ebase
no free parameters; computed column-chunked, scores identical to the library
PCA
linear projection
Ebase
components explaining 90 percent of the variance; score = sum of the distances to the retained eigenvectors divided by their explained-variance ratios
Table 4: Detectors, the supervised reference and DeepSAD with their settings, fixed in advance and never tuned on the labels; higher scores mean more anomalous. IF to VAE are PyOD implementations ( Zhao et al., 2019 ) , AE-emb and DeepSAD are our PyTorch code. Ebase is the one-hot and standardised encoding, Eemb the integer-coded input of the embedding autoencoder.
Client ledgers (five seeds)
Public
Detector
C
A
B
D
ledger
(a) P@100
IF
0.04 (0.01)
0.03 (0.02)
0.04 (0.04)
0.11 (0.07)
0.08 (0.04)
LOF
0.06 (0.03)
0.01 (0.01)
0.00 (0.00)
0.11 (0.03)
0.11
kNN
0.40 (0.02)
0.29 (0.04)
0.42 (0.03)
0.34 (0.04)
0.17
HBOS
0.03 (0.01)
0.02 (0.02)
0.03 (0.03)
0.04 (0.01)
0.04
Table 5: Conventional, hit-based view: P@100 and average precision (AP), mean (standard deviation) over seeds. Client ledgers: core grid, five seeds; supervised reference as a mean only. Public ledger: three seeds for IF, AE, VAE, AE-emb and the supervised reference, one run for the deterministic detectors. Bold: best unsupervised detector per column. The supervised LightGBM reference uses the labels and is shown for comparison only.
Figure 2: Recall within the first 100 postings per injected type, Rt @100, on the four client ledgers (core grid, mean over five seeds): (a) ledger C, (b) ledger A, (c) ledger B, (d) ledger D. Column headers give the type and its number nt of injected rows. The framed bottom row is the budget ceiling min(1,100/nt) .
Figure 3: Hits against fair-share type recall per detector: P@ k on the horizontal and FSR@ k on the vertical axis; the grey diagonal marks FSR = P. (a) to (d) Ledgers C, A, B and D at k=100 (seed means; on ledger D, bars show the rounding bounds). (e), (f) The public ledger at k=100 and k=1,000 . Axis ranges differ between panels.
Best by hits
Best by FSR@ k
Ledger ( k )
Detector
P@ k
TC@ k
FSR@ k
Detector
FSR@ k
τb
C (100)
kNN
0.40
n/d
0.57
VAE
0.65
0.83
A (100)
kNN
0.29
n/d
0.29
AE
0.38
0.65
B (100)
kNN
0.42
n/d
0.31–0.32
kNN
0.31–0.32
0.97–0.99
D (100)
PCA, AE, VAE
0.98
n/d
0.200–0.226
kNN
0.284–0.344
0.598–0.625
Public (100)
AE-emb
0.40
2.7
0.22
AE-emb
0.22
1.00
Table 6: What changes under the type-aware view. Per ledger: the detector with the highest P@ k , its type coverage TC@ k and its FSR@ k ; the detector with the highest FSR@ k ; and Kendall’s τb between the orderings of the nine detectors by P@ k and by FSR@ k . Client ledgers: FSR of the seed-averaged type profile, with ranges given by the bounds that the two-decimal rounding implies; public ledger: mean of the per-seed values.
Leader by
Kendall’s τb
Ledger ( k )
Cap ct
P@ k
Macro
FSR@ k
P, macro
macro, FSR
C (100)
nt
kNN
VAE
VAE
0.83
1.00
A (100)
nt
kNN
AE
AE
0.65
1.00
B (100)
20
kNN
VAE, AE
kNN
0.74
0.69
D (100)
20
PCA, AE, VAE
PCA, AE, VAE
kNN
0.81
0.42
Public (100)
12.5
AE-emb
AE-emb
AE-emb
0.94
0.94
Table 7: Which conclusions need the fair-share normalisation: the leader under P@ k , macro-averaged recall (the mean of Rt @ k over the types) and FSR@ k , the fair-share caps ( nt when every cap equals the type size, in which case FSR equals macro recall, otherwise the common cap k/T ), and Kendall’s τb between the orderings of the nine detectors. Client ledgers at k=100 , from the printed seed-mean per-type recalls (FSR of the seed-averaged type profile, midpoints of the rounding bounds; detectors within 0.005 of the macro-recall leader are listed together). Public ledger exact, mean over seeds.
Figure 4: The review budget on the public ledger (102,429 journal entries, 850 anomalous). (a) P@ k and (b) FSR@ k against the budget k (log scale); the dotted line in (a) is the attainable maximum and the vertical line marks k=1,000 . (c) First-hit rank per marking (dark: early; bold: within 1,000); All is the rank at which every marking has appeared.
Figure 5: Markings and rule populations on the public ledger. (a) Recall per marking within the first 1,000 journal entries; nt is the number of anomalous entries per marking. (b) P@100 when the entries inside each journal entry testing rule population are ranked by the detector score; the black bar marks the rule precision, the share of anomalies in the population.
Figure 6: Contamination and difficulty on ledgers C, A and B. (a) P@100 of kNN, VAE, AE, PCA and IF at contamination 0.5, 1.5 and 3 percent and three difficulty levels (mean over five seeds); the dotted line is the attainable maximum. (b) Exploratory critical-difference diagram of mean ranks over the 27 seed-averaged sweep cells, which derive from three ledgers and are not independent datasets; bars join detectors within CD = 2.31.
(a) Anomalies found in 100 reviews
DeepSAD, after r reviews
AE-emb, first 100
Most hits, first 100
Ledger
25
50
75
100
Found
Gain
Detector
Found
C
3.0
21.0
26.2
29.4 (4.6)
12 (2)
17.4
kNN
40
A
8.8
29.0
40.6
47.0 (3.0)
19 (1)
28.0
kNN
29
B
8.0
23.4
39.4
57.6 (11.7)
23 (7)
34.6
kNN
42
D
10.2
29.4
52.8
76.0 (8.9)
27 (11)
49.0
PCA, AE, VAE
98
Table 8: Auditor reviews fed back to DeepSAD. (a) Anomalies found after 25, 50, 75 and 100 reviews (four rounds of 25), mean over seeds with the standard deviation of the final count, against the anomalies among the first 100 units of AE-emb, the list the same reviews follow if the ranking stays static, and of the static detector with the most hits; the gain is that of the adaptive review protocol over the static AE-emb ranking. Client ledgers: core grid, five seeds; the static counts are 100 times the printed P@100 and exact to within 0.5. Public ledger: three seeds, units are journal entries. Public, control: the same protocol with the revealed labels withheld from the retraining, added in revision outside the frozen protocol. (b) Public ledger: markings among the 100 reviewed entries, with and without labels, and among the first 100 entries of AE-emb, mean over three seeds, with type coverage TC@100 and fair-share type recall FSR@100; nt is the number of entries per marking.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
N , yi
number of units; label of unit i (1 anomalous, 0 normal)
T , T
set of anomaly types (T1 to T5, or markings M1 to M8) and its size
It , nt
anomalies of type t and their number
si , π(r)
score of unit i ; unit at rank r (descending score, ties by position)
k , ρ
review budget; budget as a share of the ledger, k=ρN
topk , ft(k)
the first k units of the ranking; number of anomalies of type t among them
Appendix
Table 1: Notation of the type-aware budget evaluation, the benchmark and the review protocol.
Seed
List
M1
M2
M3
M4
M5
M6
M7
M8
F@25
F@100
TC@100
FSR@100
0
AE-emb, first 100
0
0
0
0
1
0
18
0
6
19
2
0.135
0
DeepSAD, reviewed
5
0
0
0
6
0
60
0
6
71
3
0.235
0
Control, reviewed
25
0
0
0
1
0
35
0
6
61
3
0.260
1
AE-emb, first 100
26
0
0
0
6
0
5
0
15
37
3
0.235
1
DeepSAD, reviewed
58
0
0
0
1
0
5
0
15
64
3
0.185
1
Control, reviewed
28
0
0
0
2
0
5
0
15
35
3
0.195
Appendix
Table 1: Public ledger, DeepSAD per seed: markings among the 100 reviewed entries with labels and in the no-feedback control (labels withheld, added in revision), and among the first 100 entries of the AE-emb model both runs start from, with the anomalies found after 25 and 100 reviews (F@25, F@100), type coverage TC@100 and fair-share type recall FSR@100. Round 1 reviews the first 25 entries of AE-emb, so the three lists agree after 25 reviews.
(a) Point estimate [95 percent interval]
Detector
P@100
FSR@100
P@1,000
FSR@1,000
TC@1,000
IF
0.08 [0.03, 0.14]
0.08 [0.03, 0.14]
0.083 [0.067, 0.100]
0.111 [0.087, 0.136]
7 [6, 7]
LOF
0.11 [0.05, 0.18]
0.11 [0.05, 0.13]
0.074 [0.059, 0.092]
0.128 [0.112, 0.145]
4 [3, 4]
kNN
0.17 [0.10, 0.26]
0.17 [0.10, 0.26]
0.215 [0.189, 0.240]
0.301 [0.275, 0.323]
6 [5, 7]
HBOS
0.04 [0.01, 0.08]
0.04 [0.01, 0.08]
0.067 [0.052, 0.083]
0.104 [0.077, 0.131]
7 [5, 7]
ECOD
0.04 [0.01, 0.08]
0.04 [0.01, 0.08]
0.083 [0.065, 0.100]
0.136 [0.110, 0.158]
6 [5, 6]
Appendix
Table 1: Public ledger, bootstrap over journal entries: 1,000 resamples of the entries with replacement, the same resamples for every detector, percentile 95 percent intervals. The stochastic detectors enter with their seed-0 scores, so the point estimates can differ from the three-seed means in Table 5 . (a) Intervals per detector. (b) Paired differences kNN minus the named detector; italic marks an interval that contains zero, and the second line gives the share of resamples in which kNN is strictly ahead.
Figure 1: Bootstrap intervals on the public ledger: 1,000 resamples of the 102,429 journal entries with replacement, the same resamples for every detector, percentile 95 percent intervals; the stochastic detectors enter with their seed-0 scores. (a) Point estimate and interval of P@100, FSR@100, P@1,000 and FSR@1,000. (b) Paired difference kNN minus each other detector; open markers are intervals that contain zero.
(a) Client ledgers, k=100
Ledger
P@100 leader
FSR leader, equal weights
Reversal fixed (of 11 schemes)
Exceptions
Random vectors with reversal fixed
C
kNN
VAE
10
T4 halved
91%
A
kNN
AE
11
none
89%
B
kNN
kNN (no reversal)
2
reversal to VAE when T1 is halved or T4 doubled
26%
D
PCA, AE, VAE
kNN
9
T1 halved, T4 doubled (undetermined)
63%
Appendix
Table 1: Sensitivity of FSR@ k to the type weights ( D ). (a) Client ledgers at k=100 : the leader by P@100, the FSR leader under equal weights, the number of the eleven structured schemes (equal weights and each type doubled or halved) in which a reversal between the P@100 leader and FSR is fixed by the rounding bounds, the schemes that are exceptions, and the share of 300 random weight vectors (each weight between half and twice its equal value) with a reversal fixed by the bounds. (b) Public ledger: the FSR leader under equal weights, the number of the seventeen structured schemes in which it stays the FSR leader, the smallest Kendall τb between the orderings by P@ k and FSR@ k over these schemes, and, over 2,000 random vectors, the share in which the equal-weight FSR leader stays the leader and the median τb .
Table S1: Core grid on the client ledgers (contamination 1.5 percent, medium difficulty): precision at budgets of 25, 50 and 100 postings, average precision and ROC-AUC, mean (standard deviation) over five seeds. Bold marks the best unsupervised detector per column within a ledger, on the printed values. The supervised reference is available as a mean for P@25, P@100 and AP only (n/a: not recorded).
Area
Deviation or implementation note
Injection
Ledger A, easy level, c=3 %: the pool holds 19 valid account pairs for the 36 T5 rows the cell needs, so 17 rows reuse a pair. The protocol permits drawing with replacement once the pool is exhausted; the reuse is recorded per cell.
Typology
The protocol called T5 a multidimensional mixed-data anomaly (Foorthuis Type VI). T5 changes two categorical fields only, so it is classified as Type V, subtype Va.
PCA
The score is PyOD’s own: the sum of the distances between the standardised row and the retained eigenvectors, each divided by its explained-variance ratio, not a literal reconstruction error. The covariance eigensolver replaces the full singular value decomposition (same components, scores agree to 4×10−4 relative or better, identical P@ k and AP).
ECOD
PyOD’s algorithm recomputed column chunk by column chunk, because the library implementation exhausts memory on ledger D; maximum difference to the library 10−14 on ledgers C, A and B. Used on every ledger.
Networks on ledger D
AE, VAE and AE-emb train with batch 256 for 30 epochs instead of batch 64 for 100; the protocol grants this for AE-emb, and a runtime probe extended it to the two PyOD networks.
Execution on ledger D
One thread per process for the linear algebra libraries and PyTorch; every detector runs in a checkpointing worker process; LOF and kNN are fitted on all 279,975 rows, the subsampling rule being implemented but never triggered. Scores are unaffected.
Appendix
Table S2: Protocol deviations and implementation notes. None changes an injection definition, a metric or a detector hyperparameter on ledgers C, A and B, and each is recorded per cell in the harness output.
Role
Client ledgers
Public ledger
Scored row
one posting
one posting line
Evaluated unit
one posting
one journal entry; its score is the maximum over its lines, and labels are constant within an entry
Account
Kontonummer
gl_account
Contra account
Gegenkonto
derived: the account of the largest-amount line on the opposite side of the same entry, 21 levels
Amount
Umsatz_mod_z , the modified z-score supplied with the data
amount ; log(1+x) and standardised in the encodings
Payment period
Zahlungszeitraum
entry lag in days, entered_date minus document_date ; used in the encodings
Appendix
Table S3: Mapping of the benchmark onto the public ledger, field by field. The protocol expects the eleven columns of the client ledgers; the table states which field of the public ledger takes each role, what had to be derived and what was added because the public ledger records timestamps. The rule attributes supplied with the dataset are excluded from every feature set, as its authors recommend. All detector settings are those of ledger D.
Detector
M1
M2
M3
M4
M5
M6
M7
M8
All
IF
1
51
626
309
46
217
3,415
774
3,415
min
1
37
327
74
12
140
2,157
661
2,157
max
1
62
1,108
654
92
279
4,486
944
4,486
LOF
239
3,116
472
11,107
192
3,098
1
6,895
11,107
kNN
43
185
824
15
70
1,001
53
4,641
4,641
HBOS
1
429
1,386
661
456
479
30
703
1,386
Appendix
Table S4: Public ledger, first-hit rank per marking: the number of journal entries an auditor reviews until the first anomaly of the marking appears. All is the rank at which every marking has appeared at least once (the largest first-hit rank). Stochastic detectors and the supervised reference: mean over three seeds, with the smallest and largest value over the seeds in italics; the mean of All is taken over the per-seed maxima.
Detector
M1
M2
M3
M4
M5
M6
M7
M8
nt
120
60
120
120
100
250
60
20
(a) Recall within the first 100 entries
IF
0.02
0.04
0.00
0.00
0.02
0.00
0.00
0.00
LOF
0.00
0.00
0.00
0.00
0.00
0.00
0.18
0.00
kNN
0.02
0.00
0.00
0.07
0.03
0.00
0.07
0.00
HBOS
0.02
0.00
0.00
0.00
0.00
0.00
0.03
0.00
Appendix
Table S5: Public ledger, recall per marking within the first 100 and the first 1,000 journal entries, mean over seeds for the stochastic detectors. Bold marks the best unsupervised detector per column (none where all are zero). M1 high amount, M2 high cash disbursement, M3 delayed entry, M4 very delayed entry, M5 non-working hours, M6 weekend, M7 shadow bank account, M8 cross-linked clearing; nt is the number of anomalous entries per marking.
Detector
T1
T2
T3
T4
T5
FSR@100
(a) Ledger C
nt
22
10
4
4
12
IF
0.12
0.16
0.05
0.00
0.00
0.07
LOF
0.22
0.00
0.00
0.30
0.03
0.11
kNN
1.00
1.00
0.15
0.10
0.58
0.57
HBOS
0.06
0.04
0.05
0.10
0.03
0.06
Appendix
Table S6: Recall within the first 100 postings per injected type ( Rt @100) on the core grid of ledgers C and A, mean over five seeds as printed in the result tables, with the fair-share type recall FSR@100 derived from them. nt is the number of injected rows of type t and the budget ceiling min(1,100/nt) the largest attainable Rt @100. Because the per-type recalls are seed means rounded to two decimals, FSR@100 is given as the tightest bounds consistent with the printed values and the printed P@100 (a single value where both bounds agree at the printed precision). Type coverage is not reported, because the seed-mean recalls do not determine it. T1 extreme amount, T2 extreme payment period, T3 unseen contra account, T4 unseen account, T5 unusual account pair.
Detector
T1
T2
T3
T4
T5
FSR@100
(c) Ledger B
nt
105
53
21
21
63
IF
0.02
0.04
0.00
0.00
0.00
0.04
LOF
0.00
0.00
0.00
0.00
0.00
0.00
kNN
0.29
0.20
0.01
0.00
0.01
0.31 to 0.32
HBOS
0.02
0.01
0.00
0.02
0.01
0.03
Appendix
Table S7: Recall within the first 100 postings per injected type, continued from Table S6 for ledgers B and D. On ledger D every type holds at least 336 injected rows, so no per-type recall can exceed 0.30.
(a) Precision at review budgets
Detector
P@25
P@50
P@100
P@1,000
IF
0.09 (0.02)
0.07 (0.03)
0.08 (0.04)
0.067 (0.024)
LOF
0.12
0.14
0.11
0.074
kNN
0.08
0.12
0.17
0.215
HBOS
0.08
0.08
0.04
0.067
ECOD
0.08
0.08
0.04
0.083
Appendix
Table S8: Public ledger, global ranking of all 102,429 journal entries, of which 850 are anomalous. Mean (standard deviation) over seeds 0 to 2 for the stochastic detectors and the supervised reference; the data are fixed, so deterministic detectors have a single run. Bold marks the best unsupervised detector per column on the printed values. The last three columns split the first 100 entries into anomalous entries, normal entries on a rare account (fewer than 1,000 lines) and other normal entries (means over seeds).
J1
J2
J3
J4
Rule
top-n amounts
high cash
delayed entry
weekend or non-working hours
Entries
2,000
402
8,207
16,493
Anomalies
196
190
310
415
Rule precision
0.098
0.473
0.038
0.025
(a) P@100 inside the population
IF
0.40
0.53
0.06
0.08
Appendix
Table S9: Public ledger, ranking inside the four journal entry testing rule populations. Each population is the entry set that one rule returns; inside it the rule attribute is constant, so the ranking cannot use it. Rule precision is the share of anomalous entries in the population, which is what plain rule matching delivers. Means over seeds for the stochastic detectors; bold marks the best unsupervised detector per column.
Client ledgers
Public
Detector
C
A
B
D
ledger
IF
14.1
14.7
0.6
9.8
5.1
LOF
0.1
0.1
1.5
1,022.6
478.2
kNN
0.0
0.1
1.5
1,029.0
467.1
HBOS
1.3
1.3
0.2
10.5
4.6
ECOD
0.1
0.1
0.4
25.1
4.8
Appendix
Table S10: Wall-clock seconds for one detector fit (2 CPU cores, 7 GB RAM, no GPU). Client ledgers: core-grid cells of the first execution of the benchmark, the one whose hardware is documented, mean over 5 seeds on ledgers C, A and B and over 3 seeds on ledger D. Public ledger: mean over the runs of a detector (three seeds for the stochastic detectors and the supervised reference), on two CPU cores. On ledger D and on the public ledger the three networks train for 30 epochs with batch 256 instead of 100 epochs with batch 64, so their times there are not comparable with ledgers C, A and B. n/a: not timed.
c=0.5 %
c=1.5 %
c=3 %
Detector
easy
medium
hard
easy
medium
hard
easy
medium
hard
(a) Ledger C
IF
0.02
0.01
0.01
0.04
0.04
0.04
0.09
0.09
0.08
LOF
0.04
0.03
0.02
0.11
0.06
0.03
0.07
0.04
0.02
kNN
0.15
0.14
0.14
0.43
0.40
0.37
0.75
0.69
0.64
HBOS
0.01
0.01
0.01
0.03
0.03
0.03
0.05
0.06
0.06
Appendix
Table S11: Contamination and difficulty sweep on ledgers C, A and B: P@100 for every combination of the contamination c and the difficulty level, mean over five seeds. The row Maximum is the attainable P@100, min(1,n/100) for n injected anomalies; at c=0.5 percent every ledger holds fewer than 100 anomalies. The 27 cells are the input of the rank analysis in Section 5.1, which uses the seed-averaged cells because per-seed values were not retained. Bold marks the highest P@100 per column within a ledger.
Journal Entry Tests (JETs) are a mandatory part of annual audits to evaluate and assess both highrisk audit areas and potential material misstatements. However, as JETs are designed to detect known patterns based on domain knowledge, the resulting lists are often very large and require substantial additional effort from the auditor. To ensure the economic efficiency of the audit, the number of false positives in JET result lists must be reduced. Especially machine learning (ML) methods represent a promising approach to improve anomaly detection in this field. In this research in progress paper, we investigate different approaches on how to combine JETs with ML-methods in a hybrid manner. We present specialized models to increase the detection performance and validity of anomaly detection results to improve audit efficiency. The experiments are based on synthetic data consisting of different normal and anomalous journal entries.
Jan Gronewald, Alexander Michael Rombach, Sebastian Stephan +1
German Research Center for Artificial Intelligence (DFKI) GmbH
Reconstruction-based anomaly detectors are accurate but opaque: a deep autoencoder flags a sample without telling a practitioner which feature ranges made it anomalous. We propose DIFFINT, an autoencoder whose latent bottleneck is structured as a set of soft, axis-aligned interval memberships learned end-to-end directly from raw numerical data, without any discretization or binarization. Each latent unit corresponds to a human-readable hyper-rectangle in feature space; an instance is encoded by how strongly it falls inside each interval relative to the other units, and its reconstruction error is the anomaly score. This keeps the power of differentiable representation learning while exposing an inspectable internal structure. We make the inductive bias precise: a certified reconstruction-error lower bound for points that fall outside every active coordinate of the learned support (with a Lipschitz-enforced decoder), and a graded, empirically verified suppression mechanism for the usual case in which only a few features are abnormal; and we provide a closed-form, label-free importance that ranks each (unit, feature) pair from quantities the model already maintains, turning trained intervals into auditable candidate constraints without ever seeing an anomaly label. On 48 ADBench benchmarks against 22 baselines under a common [-1, 1]-normalized protocol, DIFFINT attains the best mean rank overall on both metrics (4.10 on ROC-AUC, 4.16 on AUPR); among inlier-only detectors it leads its regime clearly, and it is competitive with the strongest contaminated-data detectors (see the stratified and complete-case analyses). It is the only interpretable detector in the statistically-tied leading cluster of seven methods.
Lamine Diop, Marc Plantevit
EPITA, Laboratoire de Recherche de l’EPITA (LRE) · Le Kremlin-Bicˆetre, France
Time-series anomaly detection trades off predictive accuracy, computational efficiency, and interpretability. We use a large language model not as the detector but as the author of one: an autonomous research loop in which the model repeatedly edits a single short NumPy program under a leakage-free objective, keeping the best-scoring detector it finds. The loop discovers two compact detectors, one for univariate and one for multivariate series, that describe short windows by their local spectral features and compare them with the training-region distribution through a covariance-aware distance. On the TSB-AD benchmark these detectors lead the field across metrics, ahead of the strongest classical, deep, and foundation-model baselines including Time-RCD, yet they train no network and use no GPU, and the multivariate detector is faster than every similarly performing baseline. LLM-driven program search is thus a practical route to accurate, efficient, and transparent detectors.
David Berghaus
Lamarr Institute Germany · Fraunhofer IAIS Germany