While deep features have transformed anomaly detection in images and video, their impact on tabular data has been less substantial, partly due to the limited availability of strong deep representations. Recently, prior-data fitted networks (PFNs) have emerged as a promising source of such representations for tabular data. In this work, we investigate how PFN representations can be adapted and leveraged for anomaly detection. The question is harder than it looks. No anomalies are available before deploy- ment, so model parameters cannot be tuned with supervision, and the reference set that defines normal behavior may itself contain the very anomalies it is supposed to reveal. We begin our study using frozen TabPFN features. Scoring each sam- ple by its distance to its nearest neighbors in feature space already gives strong results. We identify which layers to use and a feature-extraction procedure suited to the task. Next, to further improve performance, we use the reference set to fine- tune the model, so that the resulting features better separate normal samples from anomalies. On the ADBench benchmark, our fine-tuning free approach (ZEN) reaches a higher mean AUROC than every baseline, and our fine-tuned method (FOCUS) improves on it further. Our approach also generalizes across PFN models.
Figures & tables
Figure 1: Overview of ZEN and FOCUS. The reference set contains undetected anomalies (orange). (1) ZEN embeds the reference set with frozen TabPFN, using only the last four of its 24 blocks (bottom slabs). Each test sample is scored by its distance to its nearest reference samples in each block’s embedding, and the four scores are averaged. (2) FOCUS fine-tunes the model so that the reference set’s embeddings contract toward the fixed center c , and averages the weights of its epoch checkpoints (two epochs) into one adapted model. The adapted embeddings are scored as in (1). The drawing is simplified; Secs. 2.2 and 2.3 give the full methods.
Method
Contaminated
Clean
Method
Contaminated
Clean
AUROC
rank
AUROC
rank
AUROC
rank
AUROC
rank
FOCUS t (ours)
80.0
5.9
85.9
4.9
PCA
72.5
8.9
76.7
9.0
ZEN t (ours)
79.7
6.0
85.2
5.6
TCCM d
72.2
9.7
80.8
7.6
IForest
76.4
7.2
79.6
8.4
DeepSVDD d
70.9
11.4
74.7
11.8
TabPFN-PL t
74.6
8.8
78.8
9.4
OCSVM
70.1
11.4
75.6
10.8
COPOD
74.4
8.4
74.5
11.3
SOD
69.1
11.7
58.7
15.6
Table 1: Main results, sorted by contaminated-regime AUROC. Each regime has two columns: the mean AUROC ( ×100 ) over the 47 ADBench datasets, and the mean rank over the regime’s 18 methods. The best value per column is in bold and the second best is underlined. d marks a deep learning method, trained per dataset; t marks a TabPFN-based method.
Figure 2: Contaminated regime: mean rank of FOCUS and of every ranked baseline over the 47 datasets, best rank on the right (we only report FOCUS here, so ranks differ from Table 1 ). FOCUS is tested against each of the 16 baselines with a Wilcoxon signed-rank test, Holm-corrected over the 16 comparisons. Bars among the baselines join methods whose pairwise differences are not significant. FOCUS has a statistically significant advantage in every comparison. Appendix C.1 details the construction.
Figure 3: Additional results, contaminated regime. (a) Mean AUROC of the deep and TabPFN-based detectors and of FOCUS over the 47 datasets. Dashed line: the strongest baseline shown (TabPFN-PL). (b) Paired AUROC margin of ZEN over each baseline. Dot: mean; bar: 95% bootstrap interval over datasets. Every margin is positive and significant (Wilcoxon p<0.05 ); the p -value is printed where it exceeds 0.001 . † uTabPFN covers 44 of 47 datasets.
Figure 4: Ablations and analysis. (a) The ablation ladder: mean AUROC of the frozen model as ZEN’s steps are added one at a time, the plain readout, then the augmented context, the feature subsets and the soft cleaning, in both regimes. (b) The disadvantage of moving from the clean to the contaminated regime. the AUROC each method loses: FOCUS on the y -axis against raw-feature k NN on the x -axis, one point per dataset. In the shaded region FOCUS loses less due to contamination. Triangles are datasets beyond the axis. (c) ZEN’s first step, the plain readout, and full ZEN on five backbones under contamination, every constant unchanged. The asterisk marks the paper’s backbone, TabPFN-3.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
n
clean: % test normals
contam.: % test samples
breastw
683
68.2
59.9
glass
214
96.0
95.1
Hepatitis
80
100.0
100.0
Ionosphere
351
93.2
87.2
Lymphography
148
99.8
99.5
Pima
768
72.0
58.1
Appendix
Table 2: The twelve datasets the official resampling affects, with the fraction of test samples whose byte-identical twin lands in the training set under that protocol (mean over seeds 0–4). Our protocol runs these datasets at their natural size, where the fraction is zero.
Dataset
n
d
anom. %
Dataset
n
d
anom. %
ALOI ‡
49,534
27
3.0
musk
3,062
166
3.2
annthyroid
7,200
6
7.4
optdigits
5,216
64
2.9
backdoor ‡
95,329
196
2.4
PageBlocks
5,393
10
9.5
breastw †
683
9
35.0
pendigits
6,870
16
2.3
campaign ‡
41,188
62
11.3
Pima †
768
8
34.9
cardio
1,831
21
9.6
satellite
6,435
36
31.6
Appendix
Table 3: The 47 ADBench classical datasets: natural size n , dimensionality d , and anomaly rate. † below 1,000 samples (affected by the official pipeline’s resampling; run at natural size here). ‡ subsampled to 10,000 samples. § ZEN and FOCUS use a fixed 200-feature subset.
Dataset
FOCUS (ours)
ZEN (ours)
IForest
TabPFN-PL t
COPOD
HBOS
DTE d
raw k NN
ECOD
CBLOF
ALOI
56.0
55.5
53.9
61.5
52.3
54.6
55.8
59.0
54.7
56.1
annthyroid
85.9
84.3
82.9
68.7
78.0
62.9
65.1
72.3
79.1
65.6
backdoor
89.1
87.6
72.9
91.7
78.5
70.3
90.2
76.6
84.2
77.4
breastw
98.1
97.7
98.9
97.7
99.4
98.6
89.9
98.4
99.1
97.1
campaign
77.4
78.6
71.2
80.6
78.7
79.3
72.2
69.1
77.5
65.9
cardio
90.0
86.9
93.1
95.0
91.9
84.6
80.3
80.8
94.2
85.4
Appendix
Table 4: Per-dataset AUROC ( ×100 , five-seed means), contaminated regime. Methods ordered by mean AUROC; ours marked, markers as in Table 1 . The 19 methods are split over two tables for width: this table holds the stronger half, Table 5 the rest, same 47 datasets.
Dataset
PCA
TCCM d
DeepSVDD d
OCSVM
SOD
uTabPFN t †
COF
LODA
LOF
ALOI
56.5
57.3
56.2
55.1
61.0
61.0
66.1
50.4
66.6
annthyroid
66.1
74.3
69.3
59.1
79.3
87.2
70.7
44.4
71.0
backdoor
81.0
90.7
78.6
83.3
69.8
41.1
72.1
76.3
76.9
breastw
94.6
91.2
91.5
87.2
94.0
76.8
44.5
99.1
46.8
campaign
74.6
73.0
70.4
67.8
70.3
50.3
56.3
57.2
57.3
cardio
95.0
69.5
89.1
93.9
67.0
56.6
60.8
88.0
67.4
Appendix
Table 5: Per-dataset AUROC ( ×100 , five-seed means), contaminated regime: the remaining methods, continuing Table 4 over the same 47 datasets. † uTabPFN covers 44 of 47 datasets.
Dataset
FOCUS (ours)
ZEN (ours)
raw k NN
LOF
CBLOF
DTE d
TCCM d
IForest
TabPFN-PL t
HBOS
ALOI
67.7
60.0
59.2
66.8
55.2
55.1
56.4
54.5
61.4
53.6
annthyroid
91.5
90.2
86.2
91.5
73.8
73.4
83.6
91.2
82.6
70.8
backdoor
98.2
97.4
94.5
93.4
83.1
92.4
92.9
76.7
86.7
65.7
breastw
99.1
99.6
99.6
79.0
99.5
98.5
98.3
99.8
99.3
99.6
campaign
77.9
79.1
70.0
64.8
68.9
74.9
75.1
73.5
82.5
80.0
cardio
93.2
91.3
96.3
94.9
95.9
96.4
95.4
95.1
94.0
87.3
Appendix
Table 6: Per-dataset AUROC ( ×100 , five-seed means), clean regime. Methods ordered by mean AUROC; ours marked, markers as in Table 1 . The 19 methods are split over two tables for width: this table holds the stronger half, Table 7 the rest, same 47 datasets.
Dataset
PCA
OCSVM
DeepSVDD d
COPOD
ECOD
uTabPFN t †
LODA
SOD
COF
ALOI
55.6
54.8
55.4
51.7
53.3
63.3
51.5
59.0
63.2
annthyroid
80.9
66.1
84.9
77.4
78.7
93.7
58.7
71.9
64.1
backdoor
69.4
84.7
58.7
78.4
84.2
36.5
66.9
69.5
73.7
breastw
99.1
99.6
99.0
99.7
99.5
6.6
99.5
86.9
56.8
campaign
77.0
69.1
72.9
78.1
77.0
53.1
55.7
65.5
52.5
cardio
96.4
97.4
93.6
92.0
93.6
87.5
91.3
49.0
53.4
Appendix
Table 7: Per-dataset AUROC ( ×100 , five-seed means), clean regime: the remaining methods, continuing Table 6 over the same 47 datasets. † uTabPFN covers 44 of 47 datasets.
Method
Contam.
Clean
AUPRC
rank
AUPRC
rank
FOCUS (ours)
41.7
5.6
78.3
4.5
ZEN (ours)
41.1
6.3
76.9
5.3
IForest
39.1
7.7
62.3
9.1
PCA
38.2
8.6
64.6
9.5
HBOS
37.1
7.9
61.7
9.8
Appendix
Table 8: AUPRC ( ×100 ) and mean per-dataset AUPRC rank over the 47 datasets, both regimes. Ranks cover the methods with values on all 47 datasets in that regime. A value appears only where the stored score vectors reproduce that method’s official AUROC (tolerance 0.25). a TCCM’s contaminated-regime score vectors were not retained, so its AUPRC there is withheld (Appendix D.2 ). † uTabPFN covers 44 datasets and is excluded from ranks.
Contaminated
Clean
Method
mean
sd of mean
median sd
mean
sd of mean
median sd
FOCUS (ours)
80.0
0.96
2.9
85.9
0.51
1.2
ZEN (ours)
79.7
0.84
2.8
85.2
0.38
1.5
IForest
76.4
0.44
2.3
79.6
0.46
1.8
TCCM d
72.2 a
–
–
80.8
0.69
2.0
TabPFN-PL t
74.6
0.45
1.8
78.8
0.17
1.3
Appendix
Table 9: Seed variability of every method in Tables 4 – 7 : the 47-dataset mean AUROC ( ×100 ), the standard deviation of that mean over the five seeds, and the median over datasets of the per-dataset standard deviation over seeds; both regimes. Markers as in Table 1 . a TCCM’s contaminated-regime per-seed values were not retained (Appendix D.2 ); the mean is Table 1 ’s. † uTabPFN covers 44 of 47 datasets.
Figure 5: Critical-difference diagrams over the 47 datasets, best mean rank on the right; bars join methods whose pairwise Wilcoxon differences are not significant after Holm correction over all pairs (153 in each regime; we use Holm-corrected pairwise tests rather than the mean-rank post-hoc test, following Benavoli et al., 2016 ). Top: contaminated regime, 18 ranked methods. Bottom: clean regime, 18 ranked methods. uTabPFN is excluded from the ranks.
Figure 6: Multi-comparison matrix of our methods against the seventeen baselines, contaminated regime, split into two column halves, in the format of Ismail-Fawaz et al. (2023) as adopted by recent large-scale benchmark studies ( Middlehurst et al., 2024 ) . Each cell gives the mean per-dataset AUROC difference (row minus column), the win / tie / loss count for the row method, and the Wilcoxon p-value as computed, before any correction, all in bold when p<0.05 ; cell color encodes the mean difference (blue: row ahead; orange: column ahead). Rows and columns are ordered by mean AUROC. † uTabPFN covers 44 of 47 datasets.
Figure 7: Regret profiles, contaminated regime, adapted from the performance profiles of Dolan and Moré (2002) with their runtime ratio replaced by the AUROC gap: the fraction of the 47 datasets on which a method is within τ AUROC points of that dataset’s best method. Mean regret in the legend; higher and further left is better. The per-dataset best is an oracle that changes identity across datasets; every method trails it somewhere, ours least. Shown: our methods and the strongest baselines.
Figure 8: The fine-tune’s gain at each step of the ladder: the adapted model minus the frozen model, mean AUROC over the 47 datasets, both regimes.
Figure 9: The isolation estimate against the reference labels, contaminated regime: AUROC of the estimate as an anomaly score for the reference samples, computed from one plain embedding, from one representation (one feature subset, one block), from the full set of features pooled over its four blocks, and from all 24 representations pooled as in ZEN; bars are standard errors.
Step
AUROC
vs previous
p
wins
Contaminated regime, frozen model (ZEN)
plain readout
74.1
–
–
–
+ augmented context
75.0
+0.9
0.130
27/47
+ feature subsets
76.9
+1.9
< 0.001
38/47
+ soft cleaning
79.7
+2.8
< 0.001
36/47
soft cleaning per representation, not pooled
78.6
-1.1
< 0.001
10/47
Appendix
Table 10: The ablation ladder of Figure 4 a in numbers: mean AUROC over the 47 datasets at each step, the change from the previous step with its Wilcoxon p -value, and the number of datasets that improve. The last step of each block is the reported method (Table 1 ’s clean FOCUS, 85.9, recomputes here as 85.8, within the stated tolerance); in the clean regime the soft cleaning is switched off ( λ=0 ), so its row changes nothing.
Figure 10: Mean AUROC over the 47 datasets, five seeds, when the readout uses one transformer block at a time (dots), for the frozen model (ZEN) and the adapted one (FOCUS); dashed lines: the paper’s four-block window, shaded.
Readout window
Contaminated
Clean
ZEN
FOCUS
ZEN
FOCUS
last block
79.8 (+0.1, 0.65)
80.0 ( ± 0.0, 0.60)
84.7 (-0.5, < 0.001)
85.9 ( ± 0.0, 0.06)
last 2 blocks
79.9 (+0.3, 0.91)
80.0 ( ± 0.0, 0.50)
84.8 (-0.4, < 0.001)
85.9 ( ± 0.0, 0.06)
last 3 blocks
79.8 (+0.2, 0.54)
80.0 ( ± 0.0, 0.49)
85.0 (-0.2, < 0.001)
85.9 ( ± 0.0, 0.13)
last 4 blocks (paper)
79.7
80.0
85.2
85.9
last 6 blocks
79.3 (-0.4, 0.22)
79.7 (-0.3, 0.37)
85.3 (+0.1, 0.04)
85.9 ( ± 0.0, 0.48)
Appendix
Table 11: The readout window: mean AUROC over the 47 datasets, five seeds, for the frozen model (ZEN) and the adapted one (FOCUS) in both regimes, with the difference from the paper’s window and its Wilcoxon p -value. Blocks are numbered 0–23; the paper reads blocks 20–23.
Contaminated
Clean
ZEN
FOCUS
ZEN
FOCUS
Neighbor count k
5
76.9 (-2.8, < 0.001)
77.9 (-2.2, 0.002)
86.0 (+0.8, 0.03)
86.2 (+0.3, 0.35)
10
78.0 (-1.7, 0.001)
78.8 (-1.2, 0.02)
85.8 (+0.6, 0.01)
86.2 (+0.3, 0.10)
20
78.8 (-0.9, 0.004)
79.5 (-0.5, 0.06)
85.5 (+0.3, 0.007)
86.1 (+0.2, 0.05)
50 (paper)
79.7
80.0
85.2
85.8
Appendix
Table 12: Sensitivity of the readout: mean AUROC over the 47 datasets, five seeds, for the frozen model (ZEN) and the adapted one (FOCUS) in both regimes, one constant varied at a time with the difference from the paper’s value and its Wilcoxon p -value. With B random subsets the trust weight pools the (B+1)×4 representations, as ZEN does with six.
Fine-tuning variant
Contaminated
Clean
Trainable blocks K (paper: 6 contaminated, 12 clean)
K=3
80.5 (-0.5, 0.91)
85.5 (-0.5, 0.14)
K=6
81.0 (paper)
85.0 (-1.0, 0.04)
K=12
78.3 (-2.6, 0.004)
86.0 (paper)
K=24
77.8 (-3.2, 0.002)
83.1 (-3.0, < 0.001)
Other constants, at the paper’s K
Appendix
Table 13: Ablations of the fine-tune, seed 0, 47 datasets, mixed GPU types: mean AUROC of FOCUS with the difference from the paper’s setting rerun under the same conditions and its Wilcoxon p -value. Every arm keeps the paper’s other constants; the last row is the frozen model on the same runs.
Context construction
ZEN, contaminated
no synthetic samples ( m=0 )
79.4 (-0.4, 0.18)
m=100
79.2 (-0.7, 0.06)
m=min(n,500) , 10 folds (paper)
79.9
m=min(n,2000)
80.6 (+0.8, 0.56)
5 folds
80.5 (+0.6, 0.18)
2 folds
80.4 (+0.5, 0.92)
Appendix
Table 14: The context of each run, contaminated regime, seed 0, 47 datasets: mean AUROC of ZEN with the difference from the paper’s setting and its Wilcoxon p -value.
Figure 11: Injected contamination on eight datasets, five seeds. (a) Mean AUROC on a test set that is identical at every rate, as anomalies are injected into the clean reference set; ZEN with and without the soft cleaning, and Isolation Forest. (b) Mean trust weight of the injected anomalies and of the normal reference samples.
Backbone
blocks
plain
ZEN
gain ( p , wins)
ZEN > raw k NN
Contaminated regime
TabPFN v2
12
67.2
75.4
+8.2 ( < 0.001, 43/47)
22/47
Mitra
12
70.9
76.3
+5.4 ( < 0.001, 37/47)
22/47
TabICL
12
71.7
78.9
+7.2 ( < 0.001, 39/47)
26/47
TabDPT
32
74.2
79.0
+4.9 ( < 0.001, 36/47)
29/47
TabPFN-3
24
74.2
79.7
+5.5 ( < 0.001, 36/47)
28/47
Appendix
Table 15: ZEN on five backbones: mean AUROC of the plain readout and of ZEN over the 47 datasets, ZEN’s gain with its Wilcoxon p -value and the number of datasets that improve, and the number of datasets on which ZEN beats raw-feature k NN. Every constant is the paper’s.
Figure 12: Plain readout and ZEN on five in-context backbones, both regimes; dashed line: raw-feature k NN; the asterisk marks the paper’s backbone, TabPFN-3. Figure 4 c is panel (a) without the raw-feature k NN line.
and the full set of features, drawn once per dataset and seed
Combination
mean of the random subsets + full set, equal weight (Eq. 4 )
Synthetic context samples
m=min(n,500) per feature subset, uniform on [−1,1] ,
Appendix
Table 16: Complete configuration of our methods. The upper block is shared by both regimes; the lower block lists the per-regime constants.
Method
median s / dataset
total s (47 datasets)
LODA
0.03
2.7
PCA
< 0.01
3.1
COPOD
0.01
3.2
ECOD
0.01
3.3
raw k NN
0.05
3.6
HBOS
< 0.01
4.5
Appendix
Table 17: Wall-clock time over the 47 contaminated-regime datasets on one NVIDIA A100-SXM4-80GB. Baselines: fit plus scoring on the seed-0 splits, single run, indicative rather than seed-averaged. Ours: median per dataset and total per seed of each stage, over the 235 runs of the official run.