Comparative review of hybrid forecasting models for short-term prediction of building thermal load
Organizations: VIVAVIS Schweiz AG - 5405 Baden, Switzerland · Department of Electrical and Computer Engineering, University of Western Macedonia, 50100 Kozani, Greece
Abstract
In this paper, a comparative review of different hybrid models for short-term forecasting of building thermal demand is carried out. Particularly, the assessment tackles the comparison of data-driven models enhanced with other state-of-the-art techniques. At the first step, the existing techniques reported in the literature are analysed. It is concluded that Metaheuristics or a data-driven model are used to identify the parameters of the basic model. The qualitative evaluation includes for each method the input and output features, main advantages and drawbacks. At the second step, an existing dataset of historical thermal demand from Scottish households, as well as historical weather forecasts are utilized to assess additionally the performance of existing hybrid methods. From the assessment of 13 hybrid methods, the Empirical Modal Decomposition - long short-term memory - Markov (EMD-LSTM-Markov) model can predict with the highest accuracy the day-ahead power pattern of heating and domestic hot water (DHW) demands. Though local power peaks are also accurately predicted, high power swells and spikes are underestimated. Other methods, such as Support Vector Machine - Simulated Annealing (SVM-SA) and Random Forest - Improved Sparrow Search Algorithm - LSTM (RF-ISSA-LSTM) predict a smooth pattern of heating and DHW demand profiles with rapid changes underestimating most power peaks.
Figures & tables
| Ref. | Inputs | Outputs |
| Cen and Lim (2024) | • Historical outputs: Electricity demand of lighting, HVAC units and plug loads for each of the 33 building zones. • Building and indoor environment: Historical indoor ambient light (IAL), indoor ambient temperature (IAT), and indoor relative humidity (RH). | • Electricity demand of lighting, HVAC units and plug loads. • Indoor ambient light. • Indoor ambient temperature. • Indoor relative humidity. |
| Dong et al. (2022) | • Weather: Solar radiation, outdoor temperature, outdoor humidity. • Historical outputs: Cooling load during the previous six hours. | • Hourly cooling load. |
| Ren et al. (2022) | • Weather: Outdoor temperature, dew point temperature, humidity, air pressure, wind speed. • Calendar: Working day type. • Historical outputs: Electricity, heating and cooling loads (previous 4 h and previous 7 h), together with the corresponding values at , and of the previous week. | • Electricity load. • Heating load. • Cooling load. |
| Liu et al. (2023a) | • Weather: Outdoor dry-bulb temperature, outdoor relative humidity, solar radiation intensity (time-shifted). • Building and indoor environment: Indoor CO 2 concentration, indoor dry-bulb temperature, indoor relative humidity (time-shifted). • Historical outputs: HVAC energy consumption (1 h lag). | • HVAC energy consumption. |
| Yan et al. (2023) | • Weather: Air pressure, dew point temperature, air temperature, wind speed, wind direction (4 h lag). • Calendar: Month, date, day of week, day type, hour (4 h lag). • Historical outputs: Electricity, cooling and heating loads (1–24 h historical values). | • Electricity load. • Cooling load. • Heating load. |
| Yan et al. (2024) | • Weather: Outdoor temperature, relative humidity, wind speed, wind direction, global horizontal radiation, direct normal radiation, diffuse horizontal radiation. • Calendar: Month, day, hour, day of week, day type. • Historical outputs: Electricity, cooling, heating and gas loads (1–24 h historical values). | • Electricity load. • Cooling load. • Heating load. • Gas load. |
| Model | Advantages | Drawbacks |
| PatchTCN–TST Cen and Lim (2024) | • Channel-independent forecasting enables scalable multivariate prediction. • Patching improves local dependency extraction and reduces attention complexity. • TCN enhances temporal feature representation. • Superior accuracy across multiple metrics and scenarios. | • Inter-channel dependencies (e.g., temperature and AC power correlation) are neglected. |
| DwdAdam–ILSTM Dong et al. (2022) | • Captures long-term dependencies while avoiding gradient issues. • Superior convergence, stability, and accuracy over SVR, BPNN, and LSTM variants. • Strong generalization with limited training samples. | • Generalization validated mainly for commercial buildings. |
| CLSTM–AR–MFFCLA Ren et al. (2022) | • Combines linear (AR) and nonlinear (LSTM/CLSTM) modeling. • Multi-dimensional feature fusion improves spatial–temporal learning. • Stable performance for peak and non-workday loads. | • Increased architectural and computational complexity. |
| PCC–GRU Liu et al. (2023a) | • PCC-based time shifting enhances feature correlation and model stability. • GRU models nonlinear dynamics and peak-shaving behavior. • Mitigates energy storage interference. | • Limited performance improvement for SVR-based models. |
| FTTrans–E–BL Yan et al. (2023) | • Feature-time attention captures temporal and global dependencies. • Pareto-based multi-task learning improves adaptability and stability. • Probabilistic forecasting quantifies uncertainty. | • Higher computational complexity despite reduced training time. |
| TE–BiSRU–MTL Yan et al. (2024) | • Transfer entropy clustering captures nonlinear multi-feature dependencies. • Temporal attention and Bi-SRU enhance accuracy and efficiency. • Multi-task learning improves generalization across horizons. | • Increased complexity due to ensemble and multi-task frameworks. |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
| Patch length | |
| Number of patches | |
| Transformer dimension | |
| Attention heads | |
| Batch size | |
| Training epochs |
| Parameter | Value |
| LSTM batch size | |
| LSTM hidden layers | |
| LSTM hidden units | |
| Training epochs | |
| Learning rate | |
| Activation function | Softplus |
| Parameter | Value |
| Batch size | |
| Training epochs | |
| Learning rate | |
| Encoder input size | Selected features |
| Encoder hidden size | |
| Encoder layers |
| Parameter | Value |
| Training batch size | |
| Training epochs | |
| GRU layers | |
| GRU hidden units | |
| Output dimension | |
| Activation function | Softplus() |
| Parameter | Value |
| Training epochs | |
| Training batch size | |
| Learning rate | |
| Input feature embedding dimension | |
| Transformer encoder layers | |
| Transformer attention heads |
| Parameter | Value |
| Copula entropy bins | |
| Relative entropy tolerance | |
| Bi-GRU hidden size | |
| Bi-GRU output size | |
| Training epochs | |
| Training batch size |
| Parameter | Value |
| Training epochs | |
| Training batch size | |
| Learning rate | |
| Wavelet family | bior2.8 |
| WTD decomposition level | |
| CNN layers |
| Parameter | Value |
| Correlation subsample | |
| Training batch size | |
| Inference batch size | |
| Training epochs | |
| Early stopping patience | |
| Learning rate |
| Parameter | Value |
| Training epochs | |
| Learning rate | |
| Training batch size | |
| Inference batch size | |
| Optimizer | Adam |
| Teacher forcing ratio |
| Parameter | Value |
| SA cycles | |
| SA trials per cycle | |
| SA subsample | |
| LinearSVR maximum iterations | |
| LinearSVR tolerance | |
| Search space - range |
| Parameter | Value |
| PCA number of components | |
| ANN hidden layers | |
| ANN hidden neurons | |
| ANN hidden activation function | ReLU |
| ANN output activation function | Softplus |
| ANN output layer | neurons |
| Parameter | Value |
| Conv1 filters | |
| Conv2 filters | |
| Parallel convolution kernels | [3, 7] |
| Residual connection | Enabled |
| Dropout rate | |
| Adaptive pooling output length |
| Parameter | Value |
| ISSA population size | |
| ISSA candidates evaluated | |
| LSTM batch size | |
| Evaluation batch size | |
| Early stopping patience | |
| RF trees |