Tabular foundation models (TFMs) outperform established machine learning models on tabular benchmarks through in-context learning. Building on this success, interest is growing in applying them to data streams, where data arrive continuously and evolve over time. On a stream, a TFM adapts by updating its context rather than its parameters, so its accuracy and cost depend on which examples it keeps and how often it rebuilds its context. We therefore present a systematic study of TFMs on data streams, covering memory management, computational cost, and stream-specific challenges such as concept drift and delayed labels. We find that TFMs achieve the highest predictive performance and that simply retaining the most recent examples is as effective as existing memory management techniques. They also recover faster than streaming learners after drift and keep the highest accuracy under label delay. This accuracy, however, comes at a high serving cost, since a nearly unchanged context is re-encoded at every prediction. These results point to architectural efficiency as the way forward for in-context stream learning.
Figures & tables
Figure 1: Approaches to adaptation in data streams, compared under a common prequential protocol measuring predictive performance and computational cost.
Stream
Type
Instances
Features
Classes
Drift
electricity ( Harries, 1999 )
real
45,312
8
2
unknown
noaa ( Elwell and Polikar, 2011 )
real
18,159
8
2
unknown
meter ( Souza et al., 2020b )
real
22,948
96
10
unknown
rialto ( Losing et al., 2016 )
real
82,250
27
10
unknown
posture_no8 ( Kaluza et al., 2010 )
real
163,477
3
10
unknown
agr_a ( Agrawal et al., 1993 )
synth.
30,000
9
2
abrupt
Table 1: Evaluation streams.
Rank
Accuracy
Balanced acc.
Cohen’s κ
κt
Time (ms/inst.)
Memory (KB/inst.)
Method
mean
med.
mean
med.
std.
mean
rank
mean
rank
mean
rank
mean
med.
mean
med.
TabFM + FIFO
1.11
1
.869
.924
.104
.829
1.33
.772
1.22
.730
1.11
167
129
4.80
1.50
TabICL + FIFO
2.56
2
.860
.906
.105
.820
2.89
.758
2.67
.712
2.56
58.4
71.4
3.73
1.81
CURE
2.89
3
.857
.907
.106
.819
2.89
.754
2.78
.704
2.89
87.3
97.9
5.94
1.96
DR-TabPFN + FIFO
5.33
5
.838
.868
.103
.794
5.67
.724
5.33
.659
5.33
599
542
5.38
1.78
TabPFN + FIFO
6.89
6
.805
.814
.122
.762
7.33
.677
7.44
.619
6.89
103
89.8
2.13
1.32
Table 2: Predictive performance and cost of all methods, averaged over streams. Rank is the per-stream accuracy rank, time uses the geometric mean, and memory is growth per instance. Best in bold , second underlined .
Figure 2: Accuracy over the synthetic streams, each with its own drift type, in stages of roughly a tenth of the stream. Dashes mark concept changes.
Figure 3: Mean accuracy against mean cost over all streams, coloured by method family. Filled marks ( ) are on the Pareto front, hollow marks ( ) are dominated, and grey ties ( ) join each memory management system to FIFO on the same backbone.
Figure 4: Instances to reach the initial plateau (left) and to regain accuracy levels after the first abrupt drift in agr_a (right). A missing bar marks a level not reached before the next drift (dashed line).
Figure 5: Decision functions of kNN and the TFMs on the same FIFO memory shortly after a gradual drift, with the misclassified share of the input area in each title.
TabICL v2
TabPFN v1
Stream
B
FIFO
CURE
Δ
ms/inst.
FIFO
DualFIFO
Δ
ms/inst.
noaa
100
.788
.787
+.001
21.9
.784
.785
−.001
28.3
250
.796
.798
−.002
21.0
.792
.785
+.006
30.8
500
.814
.813
+.000
24.0
.808
.808
+.000
41.0
1,000
.818
.819
−.001
71.8
.814
.814
+.001
99.1
agr_a
100
.921
.918
+.003
17.3
.901
.902
−.001
24.9
Table 3: Accuracy and FIFO latency across memory budgets on a real and a synthetic abrupt-drift stream. Δ is FIFO minus the paired memory management policy. Best per stream in bold.
Figure 6: Accuracy of all methods as the label delay d grows, where each label arrives d instances after its prediction.
Context updates
Accuracy
Time (ms/inst.)
Stream
1
10
100
1,000
1
10
100
1,000
1
10
100
1,000
noaa
18,059
1,806
181
19
.818
.816
.814
.808
31.5
27.8
27.3
27.1
meter
22,848
2,285
229
23
.906
.903
.875
.710
70.3
62.0
61.3
60.7
agr_a
29,833
2,984
299
30
.935
.934
.932
.917
32.4
28.5
28.0
27.9
tree_r
59,900
5,990
599
60
.767
.768
.764
.737
35.0
31.0
30.7
30.6
Table 4: TabICL with FIFO when the context is refreshed every K instances but queried on every instance. Column headers give K , and the best per row and metric is in bold.
Time (ms/inst.)
Accuracy
Backbone
Stream
q=1
q=10
q=100
q=1
q=10
q=100
TabICL v2
meter
88.17
13.08
1.89
.906
.903
.875
tree_r
110.66
7.69
0.92
.767
.768
.764
TabPFN v1
meter
371.75
50.47
5.58
.736
.735
.709
tree_r
246.53
11.75
1.52
.646
.646
.644
DR-TabPFN
meter
1367.64
570.62
55.54
.868
.866
.835
Table 5: Query batching with q queries sharing one context, for each backbone with FIFO. Best per row and metric in bold.
Figure 7: Repeated context encoding versus cached context reuse.
TabICL v2
TabPFN v1
TabFM
Fixed cost A (ms)
26.1
59.4
120
Per-query cost γ ( μ s)
13.2
22.3
117
Queries per unit of fixed cost A/γ
1,980
2,670
1,030
Fixed share of a single call
>99.9%
>99.9%
>99.9%
Table 6: Fixed cost A and per-query cost γ fitted to t(q)=A+qγ on a separate noaa sweep. The lower block is derived from the fit.
Figure 8: Per-instance serving time (left) and peak GPU memory (right) of one compact checkpoint served from a ring cache or by re-encoding every instance, as the memory window grows. The cached path stays flat, whereas re-encoding grows in both time and memory.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Streams
Where
Default evaluation settings
Warm-up
100 instances
9
all
Evaluation window
1,000, tumbling
9
all
Label delay
0 (immediate)
9
all
Retained memory
1,000 rows
9
all
Queries per prediction
1
9
all
Appendix
Table 7: Default evaluation settings and study-specific variations.
Component
CPU host
GPU hosts
Processor
2 × AMD EPYC 7702, 128 cores
AMD EPYC 7543, 32 cores
Memory
1 TiB DDR4
251 GiB DDR4
Accelerator
none
NVIDIA RTX A6000, 48 GB
OS
Ubuntu 24.04.4 LTS
Rocky Linux 9.5
Kernel
Linux 6.8.0
Linux 5.14.0
Python
3.14
3.12 (3.11 for the v1 environment)
Appendix
Table 8: Hardware and software used for evaluation.
Method
elec
noaa
meter
rialto
posture
agr_a
sea_g
rbf_m
tree_r
Mean
Rank
No-Change
.853
.680
.011
.000
.203
.647
.561
.260
.505
.413
15.56
Majority-Class
.575
.686
.011
.100
.333
.632
.673
.343
.523
.431
15.22
HAT
.839
.744
.550
.412
.532
.883
.882
.709
.698
.694
12.33
ARF
.900
.798
.683
.728
.619
.875
.890
.888
.738
.791
8.00
SRP
.894
.794
.717
.802
.608
.914
.878
.874
.743
.803
8.11
Lev. Bagging
.890
.785
.618
.612
.600
.821
.886
.861
.691
.752
10.33
Appendix
Table 9: Per-stream accuracy with mean and mean rank across methods, lower rank is better and best values are bold.
noaa
meter
agr_a
tree_r
Mean
Rank
Method
0
100
1000
0
100
1000
0
100
1000
0
100
1000
0
100
1000
0
100
1000
No-Change
.680
.562
.557
.011
.102
.101
.647
.642
.633
.505
.502
.498
.461
.452
.447
15.8
15.5
15.8
Majority-Class
.686
.686
.686
.011
.100
.101
.632
.632
.632
.523
.523
.523
.463
.485
.486
15.2
15.5
15.2
HAT
.744
.740
.724
.550
.535
.467
.883
.878
.835
.698
.693
.646
.719
.711
.668
11.8
11.8
10.8
ARF
.798
.792
.783
.683
.632
.419
.875
.871
.831
.738
.731
.678
.773
.756
.678
8.8
8.8
9.5
SRP
.794
.791
.782
.717
.668
.446
.914
.907
.858
.743
.737
.682
.792
.776
.692
7.0
7.0
8.0
Appendix
Table 10: Accuracy under label delay d on four streams, all sixteen methods. Rank is the mean per-stream rank. Best per column in bold.
Figure 9: Temporal κt across evaluation streams, with methods ordered by median. Fill colour indicates per-instance serving cost on a logarithmic scale.
Figure 10: Mean accuracy ranks with Holm-corrected two-sided Wilcoxon comparisons against the top-ranked method.
Figure 11: Decision functions of kNN and the TFMs on the same FIFO memory after an abrupt drift, a recurrence and a gradual drift, with learned (red dashed) and true (black) boundaries and memory contents as dots.
Figure 12: Decision functions after the first abrupt drift for increasing FIFO memory sizes. Large memories still hold the previous concept.
B=50
100
200
500
1,000
slope
kNN ( k=5 )
50
50
55
130
271
0.99
TabPFN v1
50
82
109
244
271
0.58
TabICL v2
60
97
203
334
601
0.67
TabFM
60
67
165
335
601
0.80
Appendix
Table 11: Instances after the first drift until windowed accuracy regains 90 percent of its pre-drift level, by FIFO memory size B , with the slope on logarithmic axes. The smallest values sit at the measurement floor.
kNN
TabPFN v1
TabICL v2
TabFM
all
+300
all
+300
all
+300
all
+300
Best fixed size
.766
.760
.822
.780
.878
.787
.886
.780
B=50
.740
.760
.808
.780
.829
.783
.811
.767
B=1,000
.718
.593
.748
.613
.811
.610
.814
.610
Window oracle
.808
.777
.848
.807
.923
.823
.928
.830
One-switch oracle
.774
.760
.821
.780
.885
.783
.878
.767
Appendix
Table 12: Mean accuracy around the first drift (all) and just after it (+300) for fixed memory sizes, oracles and selectors.
Method
elec
noaa
meter
rialto
posture
agr_a
sea_g
rbf_m
tree_r
Mean
Rank
No-Change
0.20
0.31
0.86
0.19
0.12
0.08
0.10
0.10
0.11
0.17
1.44
Majority-Class
0.30
0.38
0.67
0.16
0.22
0.11
0.09
0.10
0.12
0.19
1.78
HAT
0.31
0.82
1.02
0.35
0.20
0.31
0.17
0.21
0.18
0.32
3.11
ARF
6.50
6.91
41.4
9.46
2.07
6.25
2.49
4.07
6.85
6.30
7.33
SRP
8.87
7.10
53.8
13.3
3.59
8.50
2.89
5.33
10.6
8.47
8.56
Lev. Bagging
2.44
2.96
35.2
6.39
1.64
2.95
1.00
2.64
2.14
3.29
5.78
Appendix
Table 13: Per-stream serving time in ms per instance and mean latency rank, with lower values indicating faster serving.
Method
elec
noaa
meter
rialto
posture
agr_a
sea_g
rbf_m
tree_r
Mean
Rank
No-Change
0.56
2.65
1.22
6.30
1.83
0.07
0.01
0.01
0.07
1.41
4.44
Majority-Class
0.74
2.77
30.3
5.97
0.18
0.18
0.02
0.05
0.05
4.48
5.11
HAT
4.88
3.79
33.6
2.16
0.80
4.27
1.02
1.28
5.27
6.34
9.89
ARF
0.63
6.87
4.83
0.56
4.41
16.6
7.11
10.3
11.1
6.93
11.22
SRP
11.9
4.88
16.1
6.35
5.57
23.4
7.15
6.46
14.8
10.7
14.00
Lev. Bagging
1.09
22.8
2.49
0.24
4.24
6.07
7.08
5.06
2.10
5.69
10.11
Appendix
Table 14: Per-stream resident memory growth in KB per instance and mean rank, with lower values indicating less growth. Dashes mark runs without a usable memory trace; means and ranks use the available streams.
Method
ms/inst.
Host
USD per 106 inst.
HAT
0.32
CPU
0.03
SAM-kNN
0.45
CPU
0.05
Lev. Bagging
3.29
CPU
0.35
OAML
3.38
CPU
0.36
ASML
4.55
CPU
0.49
ARF
6.30
CPU
0.67
Appendix
Table 15: Serving cost at on-demand cloud prices, derived from Table 13 . Prices are the AWS m6i.2xlarge rate and the median RTX A6000 rate on 9 September 2026 and are indicative only.
Windows
Accuracy
κt
Stream
prefix
tuning
scored
pretrained
updated
Δ
pretrained
updated
Δ
electricity
10,000
10
36
.9394
.9388
−.0006
+.5639
+.5600
−.0039
agr_a
7,500
8
22
.9404
.9406
+.0002
+.6570
+.6566
−.0004
Appendix
Table 16: TabICL performance before and after prefix-based fine-tuning with FIFO memory.
Stream
raw features
TabICL row repr.
TabFM + FIFO
rbf_m
.924
.720
.938
sea_g
.862
.783
.894
agr_a
.838
.805
.942
electricity
.802
.774
.950
noaa
.745
.705
.820
rialto
.618
.539
.924
Appendix
Table 17: Accuracy and serving time for similarity-based prediction using raw features, frozen TabICL representations, and TabFM+FIFO.
recency decay
drift gate
Stream
off
on
Δ
off
Δ
Real streams
rialto
.540
.618
−.078
.659
+.042
meter
.503
.527
−.024
.534
+.006
electricity
.789
.802
−.013
.815
+.013
noaa
.743
.745
−.002
.744
−.001
Appendix
Table 18: Ablation of recency decay and the drift gate in the raw-feature similarity predictor. Δ is off minus on, so negative values favour the component.
Figure 13: Datapoint-axis attention masks for bidirectional, causal, and sliding attention.
Figure 14: Cache dependencies in the evaluated TabICL-style architecture.
Figure 15: Accuracy sensitivity to the sliding-attention window size.
nanoTabPFN, 356 k
nanoTabICL, 8.1 M
Drift type
bidirectional
causal
sliding
bidirectional
causal
sliding
abrupt
.715
.880
.891
.749
.793
.784
gradual
.687
.876
.889
.743
.758
.755
incremental
.648
.819
.856
.761
.781
.780
recurring
.679
.620
.623
.737
.747
.747
stationary
.883
.875
.900
.875
.874
.870
Appendix
Table 19: Accuracy across datapoint-axis attention masks and drift conditions.
Figure 16: Accuracy across bidirectional, causal, and sliding attention masks.
W
abrupt
gradual
incremental
recurring
stationary
mean
32
.883
.896
.863
.748
.869
.852
64
.901
.913
.877
.750
.887
.866
128
.904
.916
.876
.732
.896
.865
256
.906
.912
.865
.770
.910
.872
512
.897
.917
.863
.776
.914
.874
bidirectional
.804
.831
.776
.789
.914
.823
Appendix
Table 20: Accuracy across sliding-attention window sizes for nanoTabPFN, trained separately from Table 19 .
electricity
noaa
sea_g
tree_r
Model
acc.
κt
acc.
κt
acc.
κt
acc.
κt
No-Change
.853
− .000
.680
− .000
.561
− .000
.505
− .000
TabICL + FIFO
.944
+.616
.818
+.431
.891
+.753
.767
+.530
sliding
.602
−1.71
.732
+.160
.814
+.575
.559
+.108
sliding, learned recency bias
.428
−2.90
.742
+.194
.732
+.389
.551
+.091
Appendix
Table 21: Transfer accuracy and temporal κt on evaluation streams. Bold marks the better compact model.
Accuracy
κt
Time (ms/inst.)
Peak GPU (MB)
Stream
B
Cached
Re-enc.
Δ
Cached
Re-enc.
Cached
Re-enc.
Cached
Re-enc.
electricity
128
.609
.395
+.214
−1.469
−2.824
8.2
5.6
26.5
31.3
250
.609
.609
.000
−1.469
−1.469
7.7
5.7
26.5
62.7
500
.609
.404
+.205
−1.469
−2.768
8.1
10.4
26.5
173
1,000
.609
.405
+.205
−1.469
−2.764
8.2
28.0
26.5
604
2,000
.609
.430
+.179
−1.469
−2.605
8.2
83.1
26.5
2,297
Appendix
Table 22: The same sliding-window checkpoint served from cache and by re-encoding on a stream prefix, per memory budget. Δ is cached minus re-encoded accuracy, time is per instance, and peak GPU memory is the allocator high-water mark, which depends only on the feature count and B . Electricity was timed on a second host of the same type.
Tabular stream learning requires predictions on sequentially arriving examples under distribution shift. While standard methods adapt by updating model states, tabular foundation models (TFMs) make predictions conditioned on a labeled context in an in-context manner, making them a natural alternative for stream learning. This shifts the challenge from how to update the model to how to manage the context. We propose a future information view that yields three practical requirements for context management: preserve recent examples, retain uncertain examples, and remove redundant examples. We instantiate these requirements as CURE (Context management via Uncertainty-aware admission and Redundancy aware Eviction), a context-managing policy with entropy-gated admission and redundancy-aware eviction. Across seven streams, CURE shows up to 27.0% relative improvement over classical stream learners, remains robust across multiple TFM backbones, and ranks first among other policy variants. Code and datasets are available at https://github.com/morcellinus/CURE-ICML-FMSD.
Jinmo Lee, Doyun Choi, Moongi Choi +1
Department of Computer Science and Engineering, Seoul National University, Seoul, Republic of Korea · The Kim Jaechul Graduate School of AI, KAIST, Daejeon, Republic of Korea.
Tabular Foundation Models, such as TabPFN, have received a large amount of recent attention due to their performance on in-context tabular machine learning tasks, which often exceeds classical baselines. However, practical deployment considerations of these models has received less attention. In this paper we investigate the memory requirements for these models. We demonstrate that employing model compression approaches can enable memory reductions of up to 7.6 with similar levels of performance, reducing deployment requirements by nearly 87%. Our work provides insight to practitioners seeking efficient deployment of these models in practical settings.
Shuting Luo, Monika Mikhail Kanaan, Cameron Gordon +2
Commonwealth Bank of Australia, Australia · Australian Institute for Machine Learning, University of Adelaide, Adelaide, Australia
Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context learning, but their deployment in high-throughput data streams remains challenging due to communication overhead and latency. We propose \textit{HINT}, a hierarchical inference framework that combines edge-based retrieval with cloud-based TFM inference. A graph-based approximate nearest neighbor memory maintained over a sliding window provides local predictions and uncertainty estimates, allowing confident samples to be processed locally while uncertain instances are selectively offloaded, together with their retrieved context, to a cloud-hosted TFM. The framework exposes an offloading threshold and a neighborhood retrieval policy that can be varied to balance predictive performance and communication cost. Experiments show \textit{HINT} consistently identifies favorable trade-offs.