Organizations: National Key Laboratory for Novel Software Technology, Nanjing University, China. · School of Intelligent Science and Technology, Nanjing University, China. · Siemens Data and AI Research, Beijing, China. · Nanjing University – Siemens Joint Research Center on Industrial AI, Suzhou, China.
The recent emergence of Time Series Foundation Models (TSFMs) has significantly advanced multi-step forecasting performance, enabling accurate predictions over extended future horizons. However, existing TSFMs often suffer from significantly inherent uncertainty, which typically manifests as derived forecast branches emerging at each time step and spreading to subsequent steps; different forecast branches often exhibit varying forecasting performance, thereby undermining the credibility of TSFM forecasts. In this paper, we propose the Slicing-Graphing-Alignment (SGA) method to quantify the uncertainty of multi-step TSFM forecasts. The proposed SGA first characterizes the topology of all potential forecast branches using a directed acyclic graph, such that the graph complexity bounds the uncertainty of multi-step forecasts, and then precisely measures the graph complexity by integrating both topological information and TSFM-inherent stochasticity. Experimental results conducted on 11 TSFMs and 27 datasets demonstrate that (i) SGA achieves the best performance when ranking predictive errors with uncertainty estimates; (ii) SGA works with a more extensive and more precise sampling coverage than those of existing UQ methods, deriving a quantification mechanism fundamentally different from those of established ones; and (iii) larger model scales of TSFMs correlate with lower uncertainty estimates of multi-step forecasts, suggesting another empirical scaling law for uncertainty quantification of multi-step TSFM forecasts.
Figures & tables
Figure 1 : Workflow of time series multi-step forecasting and horizon uncertainty quantification.
UQ
Vocabulary-based
Quantile-based
Trajectory-based
C-T5-Tiny
C-T5-Mini
C-T5-Small
C-T5-Base
C-T5-Large
C-2-Small
C-2
TimesFM-2.5
Timer-S1
Sundial
Aurora
Rnd
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
100.00 0.00 0
NC
93.72 11.01
100.31 9.74 0
94.93 10.75
95.22 11.09
97.99 11.81
88.23 10.79
89.81 11.59
93.94 12.00
96.99 10.70
93.93 11.37
100.31 10.35
Ppl
89.33 11.52
87.81 10.24
92.21 12.82
88.88 12.16
93.43 11.42
86.91 10.55
84.74 10.41
87.03 11.81
93.11 9.96 0
93.46 11.14
100.14 10.89
PE
84.20 11.62
84.61 9.14 0
90.89 12.66
89.47 10.44
91.77 11.86
91.21 10.53
88.80 11.21
88.10 11.64
93.91 10.14
96.07 11.91
101.83 10.77
Eig
96.73 11.65
99.70 10.54
95.01 10.73
94.55 10.76
97.62 11.35
90.93 10.77
91.37 11.27
99.21 11.88
96.23 10.33
98.10 11.94
102.20 11.10
Table 1 : Comparisons of the overall NEAURC of SGA and its contenders, averaged across 27 datasets and 11 TSFMs of 3 types, where bold and underlined values denote the best and second-best results, respectively.
Figure 2 : Visualized comparisons of the sampling coverage between SGA and NC across 11 TSFMs of 3 types.
Type
TSFM
ANCASGA∩ANC
ASGAASGA∩ANC
Vocabulary-based
C-T5-Tiny
100.00
50.24
C-T5-Mini
100.00
49.47
C-T5-Small
100.00
48.68
C-T5-Base
100.00
48.75
C-T5-Large
100.00
48.77
Quantile-based
C-2-Small
91.40
93.51
Table 2 : Comparisons of the overall overlap area ratios for 11 TSFMs of 3 types, averaged across 27 datasets.
Figure 3 : Plots of averaged uncertainty versus ARMASE of TSFMs over diverse scales, averaged across 27 datasets.
Figure 4 : Ablation comparison of the overall performance of SGA on 11 TSFMs of 3 types, averaged across 27 datasets.
Figure 5 : Impact of the number of samples K (left), the slicing length ls (middle), and the threshold coefficient λ (right) on the overall performance of SGA for 11 TSFMs of 3 types, averaged across 27 datasets.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Abbr.
Full Name
Multiple
Exploiting Inherent
Literature
Sampling?
Stochasticity?
Information-based
Ppl
Perplexity
×
✓
Fomicheva et al. (2020)
PE
Predictive Entropy
✓
✓
Malinin and Gales (2021)
Diversity-based
SE
Semantic Entropy
✓
✓
Farquhar et al. (2024)
SAR
Shifting Attention to Relevance
✓
✓
Duan et al. (2024)
SD
Semantic Density
✓
✓
Qiu and Miikkulainen (2024)
Appendix
Table 3 : Overview of the UQ contenders for LLMs.
UQ for LLMs
HUQ for TSFMs
Token sequence
Forecast trajectory
Token
Forecast value
Token-level distribution
Time-step-level distribution
Semantic equivalence
DTW-based equivalence
Semantic similarity
DTW-based similarity
Appendix
Table 4 : Correspondence between the key elements used in UQ for LLMs and HUQ for TSFMs.
Figure 6 : Overall evaluation ranking of 10 UQ methods across 27 datasets and 11 TSFMs.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
0 88.70 52.20
0 91.40 62.00
0 77.40 51.60
121.70 51.00
107.00 41.90
121.70 50.50
0 91.40 61.50
105.10 52.50
0 77.40 49.90
0 26.80 43.30
CIF 2016
138.30 21.80
176.40 38.10
136.50 30.10
132.40 17.50
126.70 20.70
132.50 17.80
169.80 37.10
128.90 19.20
135.30 30.20
0 50.10 11.40
Car Parts
118.00 4.40 0
101.80 3.70 0
121.20 4.60 0
137.60 6.00 0
142.20 6.30 0
137.70 5.70 0
129.50 5.80 0
123.70 4.50 0
121.20 4.20 0
0 57.10 1.70 0
Covid Deaths
0 76.20 13.30
0 59.10 10.70
0 79.80 17.90
0 89.70 15.70
0 95.60 15.60
0 87.60 16.30
0 72.20 8.00 0
0 93.00 17.00
0 83.70 17.00
0 56.70 12.70
Dominick
0 32.90 0.90 0
0 65.20 0.90 0
0 33.90 0.80 0
0 35.90 1.00 0
0 35.90 1.00 0
0 34.50 0.90 0
0 35.40 1.00 0
0 32.70 0.80 0
0 33.80 0.80 0
0 29.80 0.90 0
ERCOT Load
128.40 32.30
117.50 32.50
0 74.10 30.50
0 91.80 46.90
0 46.90 34.30
0 91.80 47.60
117.50 34.10
0 73.10 41.00
0 74.10 30.70
0 75.60 28.90
Appendix
Table 5 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Chronos-T5-Tiny, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
182.70 38.60
106.50 49.40
0 9.10 11.70
187.40 27.30
0 76.70 43.90
187.40 20.80
106.50 69.00
187.40 25.30
0 9.10 11.00
118.00 59.10
CIF 2016
131.80 16.30
111.40 15.00
128.20 26.50
119.60 15.40
100.80 10.20
119.70 16.00
111.20 13.90
121.40 16.40
129.90 27.80
0 37.70 6.60 0
Car Parts
120.10 4.30 0
114.40 3.90 0
130.20 5.00 0
146.70 6.30 0
139.40 6.50 0
146.40 6.30 0
137.40 6.00 0
131.90 5.20 0
127.40 4.20 0
0 57.30 1.40 0
Covid Deaths
0 87.30 12.90
0 59.70 11.50
0 69.30 13.70
0 81.70 12.80
0 81.60 12.80
0 81.00 12.90
0 60.20 36.80
0 77.60 13.20
0 75.60 14.30
0 49.70 8.40 0
Dominick
0 32.70 0.90 0
0 63.50 1.00 0
0 37.70 1.20 0
0 38.20 1.00 0
0 37.80 1.00 0
0 37.40 1.00 0
0 34.70 1.00 0
0 37.00 1.20 0
0 37.90 1.20 0
0 31.00 1.00 0
ERCOT Load
100.00 38.50
0 87.10 38.70
0 91.80 42.00
117.90 40.50
0 58.10 33.30
117.90 38.80
0 74.70 35.10
0 93.60 43.60
0 91.80 42.50
0 25.20 11.80
Appendix
Table 6 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Chronos-T5-Mini, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
0 69.00 41.20
0 71.70 48.60
0 85.40 43.40
0 52.30 35.90
0 95.60 46.60
0 52.30 36.60
0 71.70 47.20
0 52.30 36.70
0 85.40 42.30
0 12.40 18.80
CIF 2016
114.50 16.40
156.90 35.10
118.70 28.70
121.00 15.90
108.90 14.50
121.00 16.40
159.20 36.90
115.30 16.20
117.80 29.10
0 26.80 5.30 0
Car Parts
125.20 5.60 0
118.00 3.90 0
128.50 5.40 0
139.70 6.30 0
132.50 6.20 0
139.50 6.30 0
134.80 6.30 0
130.30 5.60 0
126.30 4.60 0
0 60.20 1.60 0
Covid Deaths
0 83.20 12.70
0 73.30 16.30
0 65.70 11.60
0 82.10 8.90 0
0 80.80 8.90 0
0 81.90 9.30 0
0 76.20 9.30 0
0 82.50 10.80
0 73.30 11.30
0 58.30 15.50
Dominick
0 33.60 0.90 0
0 62.70 1.00 0
0 37.90 1.10 0
0 39.10 1.00 0
0 36.90 1.00 0
0 36.70 1.00 0
0 35.30 1.00 0
0 37.40 1.20 0
0 38.30 1.10 0
0 31.20 1.10 0
ERCOT Load
119.60 39.60
0 90.80 39.00
101.60 36.10
132.80 40.40
121.60 30.20
132.80 39.40
0 96.30 43.10
124.20 39.10
101.60 36.00
0 21.30 15.80
Appendix
Table 7 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Chronos-T5-Small, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
0 89.60 54.50
118.40 54.30
119.70 42.10
0 68.10 46.70
148.50 46.00
0 68.10 46.30
118.40 54.20
0 83.20 53.70
119.70 42.40
111.90 44.10
CIF 2016
117.70 16.30
109.80 26.20
101.00 21.20
109.00 15.60
0 93.00 13.50
109.10 16.00
111.10 27.30
103.80 16.90
101.50 21.60
0 33.70 6.30 0
Car Parts
132.80 5.10 0
119.30 3.70 0
130.20 5.00 0
141.30 6.70 0
133.30 6.80 0
141.70 6.70 0
140.90 6.40 0
132.80 5.50 0
132.90 5.10 0
0 67.30 1.90 0
Covid Deaths
0 91.90 10.00
0 64.00 10.60
0 75.10 10.50
0 97.70 11.00
0 96.40 12.20
0 95.10 10.80
0 69.30 25.00
0 86.40 11.20
0 79.70 10.10
0 53.20 8.50 0
Dominick
0 33.20 0.90 0
0 59.30 1.00 0
0 41.00 1.30 0
0 40.00 1.10 0
0 37.80 1.00 0
0 38.10 1.00 0
0 35.40 1.00 0
0 39.90 1.30 0
0 41.40 1.30 0
0 31.60 1.00 0
ERCOT Load
110.50 40.40
0 98.70 41.80
0 64.20 22.70
109.50 39.30
0 64.40 31.20
109.50 38.20
0 98.70 42.70
100.30 40.90
0 64.20 22.90
0 26.80 13.50
Appendix
Table 8 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Chronos-T5-Base, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
100.50 45.70
142.80 44.10
116.90 56.10
0 85.10 47.20
0 89.60 48.50
0 85.10 47.20
142.80 42.70
0 85.10 46.00
116.90 55.90
0 44.90 51.60
CIF 2016
121.30 17.10
176.10 36.40
115.30 24.70
107.60 15.00
136.30 21.10
107.60 15.40
176.40 37.90
103.90 15.90
114.90 26.20
0 36.90 6.60 0
Car Parts
126.00 5.30 0
113.50 3.60 0
130.40 4.70 0
142.40 6.40 0
140.10 7.00 0
142.60 6.50 0
136.50 5.90 0
132.00 4.70 0
126.90 4.10 0
0 60.90 1.50 0
Covid Deaths
102.80 12.70
0 81.10 16.60
0 93.50 17.40
102.30 14.30
100.60 13.20
102.90 14.90
0 81.40 19.50
0 99.30 15.30
0 87.10 13.70
0 68.20 10.70
Dominick
0 32.20 0.90 0
0 59.90 1.00 0
0 39.00 1.30 0
0 37.90 1.00 0
0 36.60 1.00 0
0 36.30 1.00 0
0 33.90 1.00 0
0 37.90 1.30 0
0 39.50 1.40 0
0 29.90 1.00 0
ERCOT Load
143.50 51.30
0 41.80 22.70
100.50 30.00
101.10 44.10
125.30 33.00
101.10 44.20
0 53.00 27.50
101.10 44.80
100.50 31.40
120.20 44.90
Appendix
Table 9 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Chronos-T5-Large, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
0 63.40 42.30
0 63.40 44.60
120.40 43.20
120.40 42.40
129.80 43.60
120.40 41.30
0 52.00 38.30
120.40 42.80
120.40 43.50
0 21.20 36.90
CIF 2016
118.40 16.50
112.60 15.50
118.40 16.50
112.60 15.20
122.50 18.30
114.70 15.80
108.60 14.60
113.70 15.60
116.80 15.90
0 27.90 7.40 0
Car Parts
129.90 6.90 0
127.20 5.60 0
136.10 5.40 0
131.20 6.20 0
121.80 5.40 0
131.80 6.40 0
136.50 6.40 0
123.80 4.60 0
122.40 4.40 0
0 54.50 1.70 0
Covid Deaths
0 69.30 9.10 0
0 56.00 7.30 0
0 71.60 9.90 0
0 63.70 7.50 0
0 63.70 7.30 0
0 63.90 7.30 0
0 62.70 7.20 0
0 67.90 9.30 0
0 70.30 9.90 0
0 19.40 3.00 0
Dominick
0 48.50 2.50 0
0 36.10 0.70 0
0 50.80 1.30 0
0 48.10 2.40 0
0 46.40 2.40 0
0 47.30 2.40 0
0 39.40 1.20 0
0 51.60 0.70 0
0 46.40 0.80 0
0 29.80 1.00 0
ERCOT Load
140.70 40.50
142.00 41.90
144.60 42.20
118.20 40.70
0 9.60 40.00
118.20 38.90
126.10 29.50
144.60 41.40
144.60 42.80
0 66.70 40.30
Appendix
Table 10 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Chronos-2-Small, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
0 66.00 54.90
0 43.40 36.10
0 65.50 52.30
0 89.30 58.10
138.90 47.80
0 89.30 57.80
0 43.40 34.40
0 89.30 58.40
0 65.50 51.70
0 12.20 28.80
CIF 2016
111.50 15.70
100.60 15.90
105.40 14.50
106.60 14.80
113.40 16.00
106.70 15.00
0 94.20 13.80
106.60 15.20
105.20 14.70
0 24.90 5.40 0
Car Parts
117.90 5.60 0
109.90 4.00 0
130.90 5.30 0
123.20 5.50 0
114.20 4.60 0
123.70 5.60 0
124.50 4.30 0
116.50 4.40 0
120.70 4.20 0
0 54.60 1.70 0
Covid Deaths
0 70.00 8.50 0
0 54.20 9.40 0
0 55.90 8.60 0
0 63.70 7.00 0
0 62.80 7.00 0
0 63.80 6.80 0
0 63.60 8.00 0
0 54.50 6.80 0
0 51.50 6.00 0
0 21.70 3.40 0
Dominick
0 47.30 2.50 0
0 33.00 0.90 0
0 38.40 1.00 0
0 47.20 2.40 0
0 45.60 2.50 0
0 46.60 2.40 0
0 39.30 1.20 0
0 42.10 0.90 0
0 37.50 0.80 0
0 26.50 1.00 0
ERCOT Load
125.10 50.40
119.30 50.10
119.30 49.70
108.00 44.10
105.60 29.10
108.00 43.90
124.10 48.50
128.60 49.60
119.30 49.80
0 6.60 4.40 0
Appendix
Table 11 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Chronos-2, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
115.50 52.80
0 67.50 63.10
0 91.20 55.70
169.20 39.90
0 84.60 55.60
169.20 39.00
0 87.00 57.00
115.50 53.70
0 91.20 55.30
0 18.30 22.30
CIF 2016
122.90 18.80
101.00 15.70
0 99.50 16.20
123.10 17.50
143.30 23.50
125.40 19.20
101.60 16.80
108.60 18.80
105.80 20.10
0 29.40 5.10 0
Car Parts
142.50 7.50 0
0 97.90 3.10 0
0 98.10 3.90 0
136.70 6.20 0
133.10 6.60 0
136.70 6.40 0
111.40 4.20 0
0 87.90 2.90 0
0 86.80 2.80 0
0 54.80 1.60 0
Covid Deaths
0 71.70 8.10 0
0 37.80 4.70 0
0 53.70 7.00 0
0 77.00 9.10 0
0 81.70 9.70 0
0 77.80 9.20 0
0 59.70 5.80 0
0 49.80 5.00 0
0 59.40 7.90 0
0 16.20 2.40 0
Dominick
0 44.70 2.00 0
0 31.10 0.70 0
0 36.00 1.00 0
0 42.40 1.80 0
0 42.60 1.70 0
0 41.90 1.70 0
0 33.00 0.80 0
0 37.30 1.10 0
0 35.20 1.10 0
0 24.00 0.80 0
ERCOT Load
0 81.30 51.50
0 98.20 58.10
0 98.20 56.20
0 98.20 57.10
0 89.40 30.10
0 98.20 57.90
0 87.90 51.70
0 98.20 57.60
0 98.20 57.20
0 64.80 43.40
Appendix
Table 12 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for TimesFM-2.5, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
0 90.00 57.40
0 81.00 53.90
0 92.40 57.60
0 92.40 58.20
133.80 57.30
0 92.40 58.00
107.60 46.30
0 92.40 58.80
0 92.40 57.70
0 8.90 14.50
CIF 2016
114.60 18.70
0 99.70 16.60
102.40 14.10
105.10 15.10
111.80 16.30
105.60 15.10
111.40 15.90
104.30 14.90
102.00 14.00
0 32.90 6.00 0
Car Parts
158.40 7.20 0
112.70 3.80 0
127.50 5.50 0
160.20 7.00 0
154.30 7.10 0
160.30 7.10 0
123.90 3.70 0
106.70 3.40 0
115.50 4.40 0
0 59.30 1.50 0
Covid Deaths
0 67.00 7.10 0
0 25.20 3.40 0
0 25.20 3.70 0
0 68.80 7.10 0
0 68.00 7.10 0
0 68.20 6.80 0
0 59.20 5.60 0
0 34.70 4.40 0
0 31.90 3.80 0
0 12.80 2.70 0
Dominick
0 47.10 2.40 0
0 32.20 1.00 0
0 35.20 1.00 0
0 48.10 2.40 0
0 45.70 2.50 0
0 47.30 2.30 0
0 37.80 1.30 0
0 37.40 1.00 0
0 36.50 1.20 0
0 27.00 1.10 0
ERCOT Load
0 90.80 39.30
125.60 37.20
0 82.20 40.30
0 66.90 34.30
0 74.00 34.60
0 66.90 34.80
117.40 48.30
0 85.60 41.70
0 82.20 39.00
0 44.00 26.10
Appendix
Table 13 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Timer-S1, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
122.00 53.60
125.00 51.60
125.00 53.00
111.50 52.30
132.90 70.40
111.50 53.00
149.40 47.10
125.00 53.10
125.00 51.80
0 72.40 43.30
CIF 2016
111.10 16.00
111.50 14.60
110.70 14.10
105.70 14.40
0 85.50 11.10
104.30 14.00
0 99.60 12.60
107.50 14.10
108.40 13.70
0 22.90 4.80 0
Car Parts
145.50 6.40 0
146.20 6.00 0
150.60 5.90 0
140.60 5.40 0
127.20 5.60 0
140.70 5.50 0
143.60 5.50 0
144.20 5.80 0
151.00 5.70 0
0 55.90 1.60 0
Covid Deaths
0 69.90 7.00 0
0 69.00 7.40 0
0 71.60 8.20 0
0 77.50 8.70 0
0 78.30 10.90
0 76.00 8.40 0
0 69.60 7.30 0
0 71.60 8.00 0
0 71.40 8.80 0
0 12.10 1.80 0
Dominick
0 51.90 2.40 0
0 51.50 2.60 0
0 52.60 2.40 0
0 56.00 2.40 0
0 52.60 2.20 0
0 53.90 2.30 0
0 52.60 2.50 0
0 54.70 2.50 0
0 53.00 2.50 0
0 25.10 0.90 0
ERCOT Load
126.30 43.50
153.30 48.30
0 93.30 40.10
119.00 46.20
109.00 35.20
119.00 46.30
114.50 45.30
121.10 45.50
0 93.30 40.40
0 98.90 32.40
Appendix
Table 14 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Sundial, where bold and underlined values denote the best and second-best results, respectively.
Dataset
NC
Ppl
PE
Eig
Ecc
Deg
SD
SAR
SE
SGA
Australian Electricity
173.40 26.00
152.40 43.60
154.60 40.40
152.40 43.10
152.40 43.00
152.40 42.60
152.40 42.30
152.40 47.70
154.60 41.50
0 26.60 25.30
CIF 2016
130.70 16.50
124.90 12.50
129.80 13.90
119.60 16.60
103.90 18.80
119.70 17.30
113.50 17.60
122.20 15.90
131.70 15.90
0 69.10 9.80 0
Car Parts
185.30 9.40 0
178.90 8.00 0
174.40 7.70 0
165.60 7.40 0
143.50 7.20 0
165.80 7.60 0
170.30 6.90 0
168.10 8.10 0
175.10 7.80 0
0 61.00 2.10 0
Covid Deaths
106.00 15.30
106.80 15.00
117.80 15.60
120.80 16.70
119.00 15.50
120.90 16.70
117.20 16.40
121.60 17.00
119.00 16.80
0 80.60 9.80 0
Dominick
0 76.00 1.70 0
0 76.50 1.60 0
0 77.10 1.60 0
0 76.70 1.40 0
0 76.10 1.40 0
0 75.50 1.40 0
0 73.00 1.50 0
0 76.10 1.60 0
0 76.90 1.60 0
0 42.90 0.30 0
ERCOT Load
0 86.80 43.10
0 91.40 44.60
0 91.40 44.40
0 95.80 46.90
0 59.60 36.40
0 95.80 46.70
0 84.70 32.40
0 95.80 46.40
0 91.40 42.90
0 53.70 23.40
Appendix
Table 15 : Comparisons of NEAURC of SGA and its contenders across 27 datasets for Aurora, where bold and underlined values denote the best and second-best results, respectively.
Figure 7 : Plots of averaged uncertainty versus MASE of TSFMs over diverse scales across 27 datasets.
Figure 8 : Ablation comparison of UQ performance of SGA on 11 TSFMs of 3 types over the first 9 datasets.
Figure 9 : Ablation comparison of UQ performance of SGA on 11 TSFMs of 3 types over the middle 9 datasets.
Figure 10 : Ablation comparison of UQ performance of SGA on 11 TSFMs of 3 types over the last 9 datasets.
Figure 11 : Impact of the number of samples K on the performance of SGA across 27 datasets and 11 TSFMs of 3 types.
Figure 12 : Impact of the slicing length ls on the performance of SGA across 27 datasets and 11 TSFMs of 3 types.
Figure 13 : Impact of the threshold coefficient λ on the performance of SGA across 27 datasets and 11 TSFMs of 3 types.
Despite the success of Time Series Foundation Models (TSFMs) on broad benchmarks, their ability to internalize basic temporal logic, especially in settings supported by exogenous covariates, remains under-examined. We introduce SimpleTimeBench, a diagnostic univariate and multivariate "unit test" suite for primitives such as monotonic trends, periodic signals and leading indicator covariates, scenarios where near-perfect forecasts should be trivial. Surprisingly, prominent multivariate TSFMs (Chronos-2, Moirai and Toto) frequently produce suboptimal zero-shot forecasts for these inputs. While fine-tuning Chronos-2 improves its behaviour on specific tasks, we show that this adaptation degrades performance on other fundamental patterns rather than enhancing its generalizable foundational capabilities. This reveals a gap between pre-training scale and basic temporal reasoning, suggesting that current TSFMs could potentially lack the inductive biases needed to capture simple predictable functions. We further demonstrate that these failures are not merely synthetic curiosities: they persist in real-world sensor forecasting, where TSFMs consistently underutilize leading indicators available in observed covariates. This inability to capture simple relationships limits the practical utility and reliability of current multivariate models.
The deployment of Time-Series Foundation Models (TSFMs) in physical sciences is hindered by a critical trade-off: while these models encode rich, universal temporal dynamics, they suffer from severe distributional misalignment when applied zero-shot to specific scientific domains, and their computational cost prohibits deployment in edge-computing sensor networks. We address a fundamental challenge: How can we extract latent structural knowledge from misaligned foundation models (FM) to train lightweight, specialized forecasters? We propose Gated Uncertainty-Aware Routing for Distillation (Guard), a novel framework that reframes multiteacher distillation as an instance-wise decision process with two adaptive mechanisms: (1) a Contextual Router that dynamically selects the most relevant teacher based on local input statistics, exploiting complementarity across diverse foundation models; and (2) an Uncertainty-Gated Temperature mechanism that acts as a "circuit-breaker," automatically attenuating distillation strength when teacher confidence diverges from domain reality. We evaluate our proposed lightweight framework on four climate-critical domains: meteorology, ecosystem carbon flux, soil moisture, and energy grids. Our method significantly reduces RMSE relative to a fixed-weight multi-teacher distillation baseline, successfully distilling knowledge from pretrained FMs (teachers) even when they exhibit suboptimal zero-shot accuracy due to distribution shift between the original and target data domains. We demonstrate that these domain-misaligned teachers can still serve as critical correctives, outperforming the globally superior FMs on 28.5% of the hardest instances. Ultimately, this enables high-precision scientific forecasting suitable for resource-constrained edge deployment. Code is available at https://github.com/RupasreeDey/GUARD-KDD2026.
Rupasree Dey, Abdul Matin, Nathan Orwick +3
Department of Computer Science Colorado State University Fort Collins, Colorado, USA · Department of Soil and Crop Sciences Colorado State University Fort Collins, Colorado, USA
Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent open-source models demonstrate the promise of TSFMs, the field lacks a comprehensive and community-accepted model evaluation framework. We see at least four major issues impeding progress on the development of such a framework. First, existing evaluation frameworks comprise benchmark forecasting tasks derived from often outdated datasets (e.g., M3), many of which lack clear metadata and overlap with the corpora used to pre-train TSFMs. Second, these frameworks evaluate models along a narrowly defined set of benchmark forecasting tasks, such as forecast horizon length or domain, but overlook core statistical properties such as non-stationarity and seasonality. Third, domain-specific models (e.g., XGBoost) are often compared unfairly, as existing frameworks do not enforce a systematic and consistent hyperparameter tuning convention for all models. Fourth, visualization tools for interpreting comparative performance are lacking. To address these issues, we introduce TempusBench, an open-source evaluation framework for TSFMs. TempusBench consists of 1) new datasets which are not included in existing TSFM pretraining corpora, 2) a set of novel benchmark tasks that go beyond existing ones, 3) a model evaluation pipeline with a standardized hyperparameter tuning protocol, and 4) a tensorboard-based visualization interface. We provide access to our code on GitHub: https://github.com/Smlcrm/TempusBench and maintain a live leaderboard at https://smlcrm.com/tempusbench.