District heating energy hubs require reliable heat load forecasts for efficient operational scheduling. Forecasting models trained on historical data may require retraining as networks evolve. Zero-shot time-series foundation models and in-context forecasting therefore offer a promising alternative: they can adapt at inference time from recent observations rather than by repeated retraining. This study systematically evaluates TabPFN-TS and Chronos-2 for probabilistic heat load forecasting in two German district heating networks and compares them with trained baselines. We assess whether TabPFN-TS, whose underlying model is pretrained entirely on synthetic tabular rather than time-series data, can capture complex district heating dynamics. We analyze covariate choice, context length, temporal resolution, and forecast horizon on selected operating weeks, evaluate the selected configuration over the full year, and assess cross-network transfer. The principal benchmark assumes perfect weather forecasts; a separate sensitivity analysis uses retrospective weather predictions. Hourly 24-hour forecasting with a 12-week rolling context and ambient temperature provides a parsimonious configuration; longer context windows do not improve accuracy. Both TSFMs outperform all trained baselines in deterministic accuracy in the full-year benchmarks. Chronos-2 achieves the best deterministic scores, with TabPFN-TS remaining close: their CVRMSE values on the main data set are 12.48% and 13.07%, respectively. Chronos-2 also achieves lower continuous ranked probability scores in both networks, with TabPFN-TS remaining close. a TSFM-based Multi-Resolution Residual-Correction Forecaster combines an hourly base forecast with short-term high-resolution corrections. Relative to direct high-resolution forecasting, it generally reduces errors in total heat demand over 12-hour periods and recorded prediction times.
Figures & tables
Figure 1: Schematic structure of the Multi-Resolution Residual-Correction Forecaster.
Year
Hellinger distance H
KSWIN drift detections
2021
0.3921
65
2022
0.1601
26
2023
0.1827
27
2024
0.1583
31
Table 1: Distributional-change indicators for the Munich network based on 15-minute heat-load observations. Hellinger distances compare each year’s heat-load distribution with the preceding year; KSWIN drift-detection counts refer to changes detected within the indicated year.
Figure 2: Fourier magnitude spectra of the full available 15-minute heat load time series: (a) full period range and (b) periods up to one day. The x-axis is shown as period rather than frequency; marked periods indicate seasonal, weekly, daily, and sub-daily components.
Figure 3: Overall CVRMSE, CRPS, and mean steady-state prediction-call time for 24-hour heat load forecasts as a function of context length for the selected summer, winter, and transitional operating weeks.
Model
Model loading (s)
Estimated first-call overhead (s)
Steady-state inference (s/forecast)
TabPFN-TS
16.49±3.45
7.47±1.23
0.947±0.021
Chronos-2
29.78±6.14
0.260±0.019
0.01809±0.00018
Table 2: Controlled TSFM timing for hourly 24-hour forecasts with a 12-week context. Values are mean ± standard deviation across five fresh processes.
Figure 4: Overall CVRMSE for 4-hour heat load forecasts at 15-minute resolution as a function of context length.
Figure 5: CVRMSE and R2 for daily 24-hour and weekly 168-hour forecasts at hourly resolution with a 12-week context window for TabPFN-TS and Chronos-2.
Figure 6: Munich predictions for the selected 24-hour forecast setup at hourly resolution with a 12-week context window.
TabPFN-TS
Chronos-2
Weather covariates
CVRMSE (%)
R2
MAE (kW)
CVRMSE (%)
R2
MAE (kW)
Amb. temp.
11.42
0.970
102.6
10.60
0.974
96.0
Amb. temp., relative humidity
11.43
0.970
102.2
10.62
0.974
96.7
Amb. temp., wind speed
11.09
0.972
98.8
10.56
0.975
96.0
Amb. temp., precipitation
11.16
0.972
100.1
10.59
0.975
96.0
Amb. temp., precip., wind sp.
11.25
0.971
99.9
10.38
0.976
94.2
Table 3: Weather feature selection for 24-hour hourly forecasts with a 12-week context window for TabPFN-TS and Chronos-2.
Setup
CVRMSE (%)
R2
MAE (kW)
Recent 12 (auto feat.)
11.42
0.970
102.6
Recent 12
11.69
0.969
106.9
Recent 6 + Relevant 6
12.69
0.963
113.0
Table 4: Context-data selection for TabPFN-TS hourly 24-hour forecasts. Context lengths are in weeks; “auto feat.” denotes automatic temporal features.
Period
Model
CRPS (kW)
Width (kW)
Rel. w. (%)
Winter
TabPFN-TS
92.95
498.9
22.1
Chronos-2
85.56
488.4
21.6
Transitional
TabPFN-TS
93.81
370.3
28.5
Chronos-2
95.20
317.7
24.5
Summer
TabPFN-TS
25.42
114.8
36.7
Chronos-2
23.78
93.9
30.0
Table 5: CRPS and 80% prediction-interval widths for daily issued forecasts at hourly resolution with a 12-week context on the selected representative weeks.
Aggregate metrics
Mean daily rank
Model
Rank shift
CVRMSE (%)
R2
MAE (kW)
CVRMSE
R2
MAE
Munich
Chronos-2
12.48 [10.88, 14.37]
0.962 [0.946, 0.972]
85.4 [75.9, 95.6]
3.04
3.01
3.11
TabPFN-TS
13.07 [11.49, 14.94]
0.959 [0.941, 0.969]
90.3 [79.9, 101.1]
3.89
3.86
3.83
TFT
14.61 [12.92, 16.61]
0.948 [0.928, 0.961]
102.1 [90.7, 114.3]
5.88
5.85
5.93
Direct LGBM
15.87 [14.25, 17.72]
0.939 [0.917, 0.953]
113.7 [99.1, 129.6]
6.49
6.42
6.49
Table 6: Full-year benchmark and transfer-validation results for hourly 24-hour forecasts in 2024. Brackets denote 95% paired seven-day block-bootstrap confidence intervals.
Figure 7: Full-year model comparison based on daily CVRMSE. The upper panels show CD diagrams and the lower panels show pairwise win rates. Each heatmap cell gives the percentage of forecast days on which the row model outperforms the column model; exact ties count as half a win.
Figure 8: Approximate PIT histograms for the principal 2024 hourly, 24-hour forecasting benchmark in Munich and Flensburg.
Year
Static CVRMSE (%)
KSWIN CVRMSE (%)
Δ CVRMSE (percentage points)
Relative reduction (%)
2021
19.96 [18.24, 22.10]
19.09 [17.27, 21.34]
−0.88 [ −1.42 , −0.28 ]
4.4 [1.4, 7.2]
2022
19.55 [17.18, 22.68]
17.62 [14.96, 21.31]
−1.93 [ −2.77 , −1.02 ]
9.9 [4.8, 14.8]
2023
22.85 [20.40, 25.60]
19.67 [17.18, 22.73]
−3.19 [ −4.08 , −2.20 ]
13.9 [9.3, 18.1]
2024
20.69 [19.00, 22.49]
16.89 [15.09, 18.93]
−3.80 [ −4.87 , −2.65 ]
18.4 [12.7, 23.4]
Table 7: Annual mean CVRMSE across eight learned AutoGluon setups in Munich. Differences are KSWIN minus static; brackets denote 95% percentile confidence intervals.
Point forecast
Probabilistic forecast
Model
Weather cov.
CVRMSE (%)
R2
MAE (kW)
CRPS (kW)
TabPFN-TS
Realized
11.89
0.974
72.2
50.9
Predicted
12.77 (+7.4%)
0.971 (-0.4%)
77.8 (+7.7%)
55.0 (+8.1%)
Chronos-2
Realized
11.41
0.976
68.5
48.3
Predicted
12.56 (+10.1%)
0.971 (-0.5%)
76.1 (+11.0%)
53.3 (+10.2%)
Table 8: Forecast quality with realized temperatures and coherent retrospective ECMWF temperature predictions for 208 matched Munich forecast starts from 7 June to 31 December 2024.
Short-term forecast
12-hour integrated heat demand
Computation
Setup
CVRMSE (%)
R2
MAE (kW)
E-CVRMSE (%)
E-bias (%)
RTF
TabPFN-TS
Base
21.63 [19.95, 23.47]
0.897 [0.880, 0.911]
125.3 [119.0, 132.1]
8.43 [7.79, 9.12]
-0.99 [-1.38, -0.60]
2.24×10−5
HFHR
21.13 [19.37, 23.08]
0.902 [0.884, 0.916]
115.4 [109.6, 121.7]
8.78 [7.92, 9.74]
-1.77 [-2.25, -1.32]
3.86×10−4
MRRC
21.15 [19.38, 23.10]
0.902 [0.884, 0.916]
114.8 [109.1, 120.9]
8.14 [7.52, 8.82]
-1.09 [-1.44, -0.74]
2.70×10−4
Chronos-2
Table 9: Munich Base, HFHR and MRRC forecast performance pooled over 2021–2024. Brackets denote 95% percentile confidence intervals.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Annual heat provided by the Munich district heating network for complete calendar years.
Ambient temperature ( ∘ C)
Supply ( ∘ C)
Regime
Min.
Mean
Max.
Mean ± SD
Dates
Winter
-8.60
-0.51
11.20
86.42±4.42
Jan. 15–21
Transitional
0.40
7.97
24.60
83.63±4.12
Apr. 22–28
Summer
16.80
23.16
32.20
84.14±3.81
Aug. 12–18
Appendix
Table S1 : Selected representative weeks from 2024 and their ambient- and supply-temperature statistics. Supply temperatures are reported as mean ± standard deviation (SD) of the 672 unfiltered 15-minute observations in each week.
Model
Configuration
CVRMSE (%)
R2
MAE (kW)
CRPS (kW)
TabPFN-TS
M → M
11.89 [9.93, 14.46]
0.974 [0.961, 0.982]
72.2 [61.7, 83.0]
50.85 [43.22, 58.62]
F → F
14.85 [12.87, 17.12]
0.960 [0.941, 0.971]
92.0 [75.3, 109.8]
64.16 [52.57, 76.38]
M → F
13.00 [11.12, 15.36]
0.969 [0.955, 0.978]
78.9 [66.7, 91.7]
55.96 [47.09, 65.16]
M+F → F+F
12.77 [10.91, 15.13]
0.971 [0.957, 0.979]
77.8 [66.1, 89.9]
54.95 [46.52, 63.67]
Chronos-2
M → M
11.41 [9.34, 14.10]
0.976 [0.963, 0.984]
68.5 [58.9, 78.4]
48.32 [41.31, 55.47]
F → F
12.74 [10.83, 15.18]
0.971 [0.956, 0.979]
77.0 [65.0, 89.6]
53.85 [45.20, 62.79]
Appendix
Table S2 : Temperature-covariate configurations on matched Munich forecast starts. The left and right sides of each arrow specify historical and future covariates.
Model
CVRMSE (%)
R2
MAE (kW)
Chronos-2
15.64 [13.44, 18.05]
0.941 [0.914, 0.958]
109.7 [93.8, 126.9]
TabPFN-TS
17.06 [14.58, 19.94]
0.929 [0.895, 0.950]
120.2 [102.2, 140.1]
Appendix
Table S3 : Full-year Munich results for non-overlapping weekly 168-hour forecasts at hourly resolution.
Accurate short-term load forecasting (STLF) is essential for the reliable and efficient operation of modern power systems. While time series foundation models (TSFMs) have recently demonstrated remarkable performance across a wide range of forecasting tasks, their effectiveness for STLF under realistic operational conditions remains largely unexplored. In this paper, we present a comprehensive benchmark of four trained-from-scratch (TFS) models and four TSFMs across three real-world load forecasting datasets under operational scenarios that differ in the availability and quality of future covariate information. Our results show that Chronos-2 consistently achieves state-of-the-art performance in both zero-shot and fine-tuned settings when future covariates are available or accurately forecast. However, its performance degrades as covariate forecasts become increasingly noisy, whereas TimesNet exhibits greater robustness under severe covariate uncertainty. These findings demonstrate the effectiveness of covariate-informed TSFMs for STLF while highlighting the critical role of robust covariate modeling in real-world forecasting applications.
Tomas Kaljevic, Ivan Arzola, Yu Zhang
Department of Electrical and Computer Engineering University of California, Santa Cruz
Accurate load forecasting at multiple grid levels is essential for future smart grids, ranging from aggregated control area forecasts for balancing supply and demand to forecasts of individual end-consumer loads for demand-side management and energy management systems. We present a comprehensive benchmark for load forecasting across grid levels, comprising three datasets that represent a transmission system operator control area, low-voltage grid feeders, and individual end consumers. We evaluate ten methods for short-term load forecasting and find that Transformer-based approaches consistently outperform established methods, reducing forecast error by 6.6-10.7 %. To analyze the impact of architectural design, we introduce YAformer, a flexible Transformer architecture that integrates modifications from prior work and is optimized via hyperparameter optimization. However, the standard Transformer achieves superior performance, suggesting that these architectural modifications are not required for accurate load forecasting. We further evaluate the Transformer-based time-series foundation model Chronos-2, which demonstrates competitive zero-shot performance on two datasets but fails to accurately capture special events in the TSO data. Detailed analyses reveal model-specific strengths and weaknesses, and ablation studies highlight the importance of long input contexts, covariates and continuous retraining - aspects that are often overlooked in the time-series forecasting literature.
Matthias Hertel, Sebastian Pütz, Jonathan Kolar +3
Karlsruhe Institute of Technology, Germany · Helmholtz AI, Germany
Obtaining an accurate short-term forecasting for heat demand is an essential part of operating district heating networks cost-efficient and reliable. Heat consumption time series at the building level are highly dependent on exogenous variables such as outdoor temperature and individual usage patterns, making forecasting in this context a challenging task. Thus, this paper benchmarks novel Transformer-based and xLSTM architectures for short-term heat-demand forecasting. Using hourly data from 25 German buildings (2017-2025), we compare three-hour and 24-hour forecasting horizons relevant for intraday control and day-ahead scheduling. We establish a multi-building benchmark that tests whether models trained on pooled, heterogeneous building data are able to generalize across diverse building stock. The results show that the xLSTM achieves the lowest RMSE (19.88 kWh for three-hour, 21.47 kWh for 24-hour forecasts), while the Temporal Fusion Transformer attains the best MAE (9.16 kWh for three-hour forecasts). As xLSTMs and Transformers require long training times and have a huge number of trainable parameters, their sustainability remains questionable. Therefore, this paper further investigates the trade-off between predictive accuracy and computational resource demand of the evaluated forecasting models. The findings indicate that also low-parameter models like a traditional fully-connected network achieve good predictive results, highlighting that marginal accuracy gains of the novel prediction models come at substantial resource expense for this use case.
Marja Wahl, Daniel R. Bayer, Sven Rausch +1
RAUSCH Technology GmbH, Würzburg, Germany · Modeling and Simulation, University of Würzburg, Würzburg, Germany