Deep Learning vs. Statistical Models for Multi-Horizon Price Forecasting of Second-Hand Electronics: A Systematic Benchmark
Authors: Mateusz Buczyński, Michał Woźniak, Konrad Kaczyński, Anna Wróblewska, Sebastian Kuk
Organizations: Faculty of Economic Sciences, University of Warsaw, Dluga 44/50, Warsaw, 00-241, PL · WeSub, Branickiego 15, Warsaw, 02-972, PL · Faculty of Mathematics and Information Science, Warsaw University of Technology, Koszykowa 75, Warsaw, 00-662, PL
Forecasting resale prices of used electronics is critical for subscription-based platforms where pricing errors translate directly into risk. Unlike structured financial markets, second-hand electronics exhibit high volatility, sparse listing histories, and non-normal price dynamics - yet no systematic time-series benchmark exists for this domain. This paper presents the first multi-horizon benchmark of statistical and deep learning forecasting models for used electronics price prediction. We use a large-scale dataset of daily price listings from Polish online marketplaces (January 2022 to March 2025, 100+ smartphone and laptop models) and evaluate eleven models across six horizons from 1 to 365 days, covering classical methods (ARIMA, ETS, Theta), recurrent and convolutional networks (LSTM, TCN), and modern deep architectures (N-BEATS, N-HiTS, TFT, PatchTST, Informer). Three complementary evaluation protocols assess trajectory fitness, one-shot endpoint accuracy, and cross-horizon transfer. N-BEATS achieves the lowest MAPE beyond 30 days, reaching 8.51% at 365 days versus 14.94% for the best statistical baseline - a 43% reduction. At short horizons (1-7 days), all models converge near 0.72% MAPE and the naive baseline remains competitive. A single N-BEATS model trained at 365 days generalizes to all shorter horizons, eliminating the need for horizon-specific models. N-BEATS and N-HiTS also demonstrate superior hyperparameter stability.
Figures & tables
Reference
Domain
Approach
Task
Gap / Limitation
Hankar et al., 2022b , Ganesh & Venkatasubbu, 2019a , Dutulescu et al.,2023a
Used cars
ML (RF, NN, XGBoost)
Tabular regression
No time dimension; point-in-time price only
Truong et al., 2020a , Wang et al.,2021a
Real estate
ML / DL + images
Tabular regression
Structured market, rich listing attributes
Ghani, 2005a , Raykhel & Ventura,2009a
Used electronics (laptops)
Statistical, KNN
Tabular regression
Small scale; no time-series evaluation
Ali et al.,2018a
Used electronics (Kaggle)
NN
Tabular regression
Single competition dataset; no temporal modelling
Fathalla et al.,2020a
Multi-category e-commerce
Multimodal DL
Tabular regression
Features per listing; no aggregate time-series
Carta et al., 2019a , Fathalla et al.,2023a
Amazon products
ARIMA
Time-series
Single model; no DL comparison; stable market
Table 1: Summary of surveyed studies on used goods price prediction and related time-series forecasting
Figure 1: A schema of the time series preparation system
Abstract item
Predicted price
Market price
apple iphone 13 pro max 512gb grafitowy
4306.70
4250
samsung galaxy a13 128gb
730.45
400
samsung galaxy s22 ultra 256gb
3288.09
3200
acer aspire a515 4gb 128gb
1203.23
1150
apple macbook pro m2 512gb space gray
5416.09
6000
Table 2: Exemplary results of pricing using our matching algorithm with the best prediction per item compared with its business value
Statistic
Value
Date range
January 2022 – March 2025
Number of calendar days (series length)
1,186
Unique abstract items
357
of which smartphones
229
of which laptops
127
of which other
1
Table 3: Quantitative characteristics of the dataset used in this study
Figure 2: Exemplary historical postprocessed time series of used goods analysed in this paper
Figure 3: Illustrative definition of temporal cross-validation utilized in this paper
Metric name
Formula
Mean Absolute Percentage Error (MAPE)
n1∑i=1nAiFi−Ai=n1∑i=1nAPEi (9)
Winsorized MAPE (W-MAPE)
n1∑i=1nwinsorize(APE,limits=[0.1,0.1]) (10)
Symmetric Mean Absolute Percentage Error (SMAPE)
n1∑i=1n∣Ai∣+∣Fi∣2∣Fi−Ai∣ (11)
Root Mean Squared Error (RMSE)
n1∑i=1n(Fi−Ai)2 (12)
Fourth Root Mean Fourth Power Difference (FRMFPD)
(n1∑i=1n(Fi−Ai)4)41 (13)
Quantile of Absolute Percentage Error ( Qα(APE) )
inf{x∈R:P(∣APE∣≤x)≥α} (14)
Table 4: Evaluation metrics utilized in the framework
Horizon
Windows in OOS
No OOS
1
30
30
7
14
9
30
12
9
90
8
3
180
8
3
365
4
3
Table 5: Number of windows in OOS and number of OOS in the experiment
Model
Trials logged
Median (s)
Max (s)
N-BEATS
7,137
14
183
N-HiTS
6,900
15
39
TFT
29,567
9
333
PatchTST
6,902
18
341
Informer
13,350
86
2,447
Table 6: Per-trial training time during hyperparameter search (wall-clock seconds on a single GPU)
Horizon
Model
MAPE
W-MAPE
SMAPE
RMSE
FRMFPD
Q90(APE)
Wasserstein (APE, U[0,5%])
1
NHITS
0.0072
0.0055
0.0072
45.7661
102.7146
0.0191
0.0183
PatchTST
0.0072
0.0055
0.0072
45.7619
102.7242
0.0191
0.0185
Naive
0.0072
0.0055
0.0072
45.7619
102.7242
0.0191
0.0194
TCN
0.0072
0.0055
0.0072
45.7617
102.7242
0.0191
0.0188
LSTM
0.0072
0.0055
0.0072
45.7615
102.7241
0.0192
0.0182
NBEATS
0.0072
0.0055
0.0072
45.9056
102.697
0.0186
0.0193
Table 7: Metric results over whole horizon for the tested models with all tested forecast horizons
Figure 4: MAPE and RMSE results for all tested models with all forecast horizons
Figure 5: Forecasts for exemplary items for forecast horizon = 1
Figure 6: Forecasts for exemplary items for forecast horizon = 365
Horizon
Model
MAPE
W-MAPE
SMAPE
RMSE
FRMFPD
Q90(APE)
Wasserstein(APE, U[0,5%])
1
NHITS
0.0072
0.0055
0.0072
45.7661
102.7146
0.0191
0.0188
PatchTST
0.0072
0.0055
0.0072
45.7619
102.7242
0.0191
0.0173
Naive
0.0072
0.0055
0.0072
45.7619
102.7242
0.0191
0.0193
TCN
0.0072
0.0055
0.0072
45.7617
102.7242
0.0191
0.0187
LSTM
0.0072
0.0055
0.0072
45.7615
102.7241
0.0192
0.0189
NBEATS
0.0072
0.0055
0.0072
45.9056
102.697
0.0186
0.0176
Table 8: Metric results for point forecast at horizon end for the tested models with all tested forecast horizons
Figure 7: MAPE and RMSE results for all tested models with all forecast horizons for one-shot comparisons
Horizon
Model
MAPE
W-MAPE
SMAPE
RMSE
FRMFPD
Q90(APE)
Wasserstein(APE, U[0,5%])
1
ETS
0.0066
0.0052
0.0067
40.9375
103.7432
0.0164
0.0181
ARIMA
0.0068
0.0053
0.0068
41.4634
110.2254
0.0157
0.0175
Theta
0.0071
0.0057
0.0071
41.5503
103.5205
0.0168
0.0174
TFT
0.0072
0.0058
0.0072
41.5557
102.996
0.017
0.0173
Naive
0.0072
0.0058
0.0072
41.5559
102.9942
0.017
0.018
NBEATS
0.0072
0.0059
0.0073
40.2122
105.1397
0.0173
0.0189
Table 9: Metric results for the tested models trained to predict 365-day horizon tested at all horizon lengths horizon end
Horizon
Model
learning_rate
batch_size
max_steps
input_size
1
NBEATS
0.006741
0.000000
630.026454
3.492214
NHITS
0.015968
0.000000
293.711877
2.366432
TCN
0.025125
4.800000
250.000000
0.000000
LSTM
0.030758
5.962848
247.767812
0.000000
TFT
0.022626
40.209231
625.610813
0.000000
PatchTST
0.023423
103.395164
2158.703314
8.411896
Table 10: Standard deviation of selected hyperparameters across models and forecast horizons
Horizon
Model
MAPE
Mean MAPE hypertuning
t-test p-value
Wilcoxon test p-value
1
NHITS
0.0072
0.0074
0.347
0.612
PatchTST
0.0072
0.0283
0.0001
0.0
TCN
0.0072
0.0283
0.0001
0.0
LSTM
0.0072
0.0283
0.0001
0.0
NBEATS
0.0072
0.0282
0.0001
0.0
TFT
0.0072
0.0283
0.0001
0.0
Table 11: Results of MAPE for final model and during hypertuning for tested deep learning model
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
h
Architecture
max_steps
LR
N-BEATS
1
stack=[ identity ], n_blocks=[4], act=SELU
2000
1.8×10−4
7
stack=[ trend ], n_blocks=[4], act=SELU
2000
3.8×10−3
30
stack=[ trend ], n_blocks=[4], act=SELU
2000
8.8×10−5
90
stack=[ trend ], n_blocks=[4], act=SELU
2000
3.6×10−4
180
stack=[ identity ], n_blocks=[1], act=ReLU
1500
3.6×10−3
365
stack=[ identity ], n_blocks=[1], act=ReLU
1500
3.6×10−3
Appendix
Table 12: Best hyperparameter configuration per model and forecast horizon, extracted from Ray Tune random search logs (lowest validation MAPE trial across all cross-validation splits); h=180 configurations are used for h=365.
Model
Parameter
Search range / candidates
All DL models
learning_rate
U[10−5,10−1] (log-uniform)
batch_size
{32,64,128,256}
max_steps
Model-specific (see below)
input_size
Model-specific (see below)
N-BEATS
max_steps
[100,2000]
input_size
[1,28] (multiplier of h )
Appendix
Table 13: Hyperparameter search ranges used in random search tuning (75 trials per model per cross-validation split). Architecture-specific parameters are listed beneath each model’s common parameters. All models had their architecture parameters searched.
Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear. We compare nine variants from five foundation model families, evaluated in zero-shot mode, with two state-of-the-art electricity price forecasting benchmarks in Germany, Poland, and Spain over 2021-2025. Their performance is assessed in terms of point and probabilistic forecasting accuracy, as well as economic value in battery energy storage arbitrage. Only the TabPFN models consistently and significantly outperform the benchmarks across all three markets and all statistical measures. However, this statistical dominance does not translate directly into economic dominance: TabPFN performs best under unlimited bids and riskier quantile-based strategies, whereas the Distributional Deep Neural Network benchmark is more profitable when risk tolerance is lower. Thus, foundation models cannot universally replace market-specific models, and their value depends on both model architecture and the decision problem.
Arkadiusz Lipiecki, Rafał Weron
Department of Computational Social Science, Wrocław University of Science and Technology, Poland · Department of Operations Research and Business Intelligence, Wrocław University of Science and Technology, Poland · Center for Research in Energy (CoRE), Aarhus University, 8000 Aarhus C, Denmark
Unlike Business-to-Consumer e-commerce platforms (e.g., Amazon), inexperienced individual sellers on Consumer-to-Consumer platforms (e.g., eBay) often face significant challenges in setting prices for their second-hand products efficiently. Therefore, numerous studies have been proposed for automating price prediction. However, most of them are based on static regression models, which suffer from poor generalization performance and fail to capture market dynamics (e.g., the price of a used iPhone decreases over time). Inspired by recent breakthroughs in Large Language Models (LLMs), we introduce LLP, the first LLM-based generative framework for second-hand product pricing. LLP first retrieves similar products to better align with the dynamic market change. Afterwards, it leverages the LLMs' nuanced understanding of key pricing information in free-form text to generate accurate price suggestions. To strengthen the LLMs' domain reasoning over retrieved products, we apply a two-stage optimization, supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO), on a dataset built via bidirectional reasoning. Moreover, LLP employs a confidence-based filtering mechanism to reject unreliable price suggestions. Extensive experiments demonstrate that LLP substantially surpasses existing methods while generalizing well to unseen categories. We have successfully deployed LLP on Xianyu\footnote{Xianyu is China's largest second-hand e-commerce platform.}, significantly outperforming the previous pricing method. Under the same 30% product coverage, it raises the static adoption rate (SAR) from 40% to 72%, and maintains a strong SAR of 47% even at 90% recall.
Hairu Wang, Sheng You, Qiheng Zhang +5
Xianyu of Alibaba, Hangzhou, China · University of Science and Technology of China, Suzhou, China
Recent advances in Time Series Foundation Models (TSFMs) promise zero-shot forecasting capabilities with minimal task-specific training. While these models have shown strong performance across generic benchmarks, their applicability in volatile, complex electricity markets remains underexplored. Addressing this gap, this study provides a systematic empirical evaluation of several TSFMs, specifically Chronos-2 and Chronos-Bolt (developed by Amazon), and TimesFM 2.5 (provided by Google), for forecasting Belgian day-ahead and imbalance electricity prices. For both considered markets, Chronos-2 in ARX mode produces the most accurate forecasts. Compared with the best ensemble prediction from other machine learning methods, Chronos-2's Mean Absolute Error (MAE) is 5% lower for the day-ahead market. In contrast, the model yields 10% higher MAE predicting imbalance prices across all forecast horizons, except for the two-hour-ahead horizon. Moreover, we find that TSFMs exhibit genuine zero-shot forecasting skills but still struggle under extreme market conditions.
Chi Bui, Maria Margarida Mascarenhas, Arnaud Verstraeten +1