The current trend in the literature on Time Series Classification is to develop increasingly accurate algorithms by combining multiple models in ensemble hybrids, representing time series in complex and expressive feature spaces, and extracting features from different representations of the same time series. As a consequence of this focus on predictive performance, the best time series classifiers are black-box models, which are not understandable from a human standpoint. Even the approaches that are regarded as interpretable, such as shapelet-based ones, rely on randomization to maintain computational efficiency. This poses challenges for interpretability, as the explanation can change from run to run. Given these limitations, we propose the Bag-Of-Receptive-Field (BORF), a fast, interpretable, and deterministic time series transform. Building upon the classical Bag-Of-Patterns, we bridge the gap between convolutional operators and discretization, enhancing the Symbolic Aggregate Approximation (SAX) with dilation and stride, which can more effectively capture temporal patterns at multiple scales. We propose an algorithmic speedup that reduces the time complexity associated with SAX-based classifiers, allowing the extension of the Bag-Of-Patterns to the more flexible Bag-Of-Receptive-Fields, represented as a sparse multivariate tensor. The empirical results from testing our proposal on more than 150 univariate and multivariate classification datasets demonstrate good accuracy and great computational efficiency compared to traditional SAX-based methods and state-of-the-art time series classifiers, while providing easy-to-understand explanations.
Figures & tables
Fig. 1: Two receptive fields (blue and red) with a length of w=9 , divided into segments of length q=3 . Dilation is the number of steps between consecutive observations within the receptive field. d⋅q is the segment hop, i.e., the number of steps between consecutive segments. The parameter s=5 is the stride, i.e., the distance between consecutive receptive fields.
Fig. 2: A simplified schema of Algorithm 2 for the two receptive fields of Figure 1 , extracted from signal xi,j . First, the moving average μseg is computed. Values in μseg are used to fill M as in Equation 11 . The two receptive fields are then normalized, binned into the SAX words [0,1,0] and [0,1,1] , and hashed into the integers and respective counts 10,1 and 11,1 . This allows the update of Z . E.g., for the red word, zi,j,11=1 .
Fig. 3: Example instances from the three classes of the “Cylinder-Bell-Funnel” dataset.
Fig. 4: Local explanation for a Cylinder instance of the CBF dataset. On the left, the original time series colored based on the saliency map, highlighting the most relevant observations for the classification. On the right, the medoid of the most important pattern that is not contained in the time series. The grey area represents all the possible alignments of the pattern ‘0,1,1,2’ found in the dataset.
Fig. 5: Global explanation for three of the most important patterns for the CBF dataset. On the left, in dark gray, the medoid of all alignments of the pattern in the CBF dataset. On the right, the count of appearances of each pattern in each time series in the dataset, divided by label. The color represents the importance of the pattern count in terms of SHAP values.
interpretable
Model
features
classifier
deterministic
Dictionary-Based
BORF
(ours)
✓
✓
✓
BOP
[ 14 ]
✓
✓
✓
HYDRA
[ 2 ]
✗
✓
✗
MRSEQL
[ 19 ]
✓*
✓
✓
TABLE I: List of competitor classifiers, divided by classifier family, with information about feature space and classifier interpretability, and stochasticity. Models that have an interpretable feature space and classifier, and are also deterministic are highlighted in gray.
Fig. 6: CD Plots for all benchmarked methods. Best models to the right.
Fig. 7: CD Plots for SAX-based, dictionary-based, and interpretable methods. Best models to the right.
Δ accuracy
median
mean
win
tie
loss
p-val
MR-HYDRA
+0.022192
+0.037400
109
22
27
0.000100
RDST
+0.020242
+0.030600
103
20
35
0.000100
ROCKET
+0.016769
+0.022100
100
17
41
0.000100
HYDRA
+0.011855
+0.022700
95
19
44
0.000100
MINIROCKET
+0.010000
+0.020500
96
17
45
0.000100
TABLE II: Comparison Matrix between BORF and competitor approaches. Models are sorted by median difference in accuracy with BORF . Wins, ties, and losses are to be read “against BORF ”. Best models on top. Grey rows are statistically tied with BORF using the Wilcoxon signed-rank test.
Δ runtime (sec)
median
mean
win
loss
p-val
BORF
1.305673
298.748806
HYDRA
+0.240600
−294.475100
55
103
0.4150
MRSQM
+1.445971
+1717.956800
16
142
0.000100
MR-HYDRA
+1.842159
−286.187900
25
133
0.000100
SAXVSM
+2.037075
+8465.654600
14
144
0.000100
TABLE III: Comparison Matrix between BORF and competitor approaches. Models are sorted by median difference in runtime with BORF . Wins, and losses are to be read “against BORF ”. Best models on top. Grey rows are statistically tied with BORF using the Wilcoxon signed-rank test.
Fig. 8: Comparison of median accuracy performance against median runtime (seconds). Best models are on the top left.
Fig. 9: Runtime comparison between naive PAA and our proposal (wPAA). On the left: runtime when changing the window size, and keeping time series length fixed. On the right: runtime when changing the time series length and setting the window size to half the number of points.
Fig. 10: Space complexity. On the left: dataset size against the number of extracted patterns. Each blue point represents the empirical number of extracted features for each dataset, and each orange point is the upper bound on the number of features for that specific dataset. On the right, the dataset size against non-zero elements in Z and fitted log-log regression line.
θ
d
ash
l
w
Acc
Time (s)
baseline
✓
✓
✓
2,4,8
all
0.820175
17.954614
w/o θ
✗
✓
✓
2,4,8
all
−0.012570
+7.954003
w/o ash
✓
✓
✗
2,4,8
all
−0.019762
−0.053850
w/o dilations
✓
✗
✓
2,4,8
all
−0.009656
−10.144182
maxl=2
✓
✓
✓
2
all
−0.185247
-17.196495
maxl=4
✓
✓
✓
2,4
all
−0.050413
-14.253021
TABLE IV: Performance delta in median Accuracy (higher is better) and median runtime (lower is better) for various alternative hyperparameter configurations with respect to the BORF baseline heuristic presented in Section V . The two best-performing models for each metric are in bold.
Fig. 11: Electrode locations based on the International 10-20 system for encephalography recording. The FingerMovements dataset contains recordings from the blue regions.
Fig. 12: Local explanation for a Left-Hand instance of the FingerMovements dataset. On the left, 12 out of the 28 signals of the original time series colored based on the saliency map, highlighting the most relevant observations for the classification. On the right, for each signal, the medoid of the most important pattern that is not contained in that signal of the time series.
Fig. 13: Global explanation for four of the most important ‘2,0’ patterns for the FingerMovements dataset. On the left, in dark gray, the medoid of all alignments of the pattern in the FingerMovements dataset. On the right, the count of appearances of each pattern in each time series in the dataset, divided by label. The color represents the importance of the pattern count in terms of SHAP values.
Figure 27Figure 28Figure 29Figure 30
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Time Series Data
X,X,x,x
time series dataset, instance, signal, entry
y,y
labels vector, value
n
number of instances in a dataset
c
number of signals in a time series
m
number of observations in a signal
Convolution
Appendix
TABLE V: Summary of notation.
Fig. 14: CD Plots for SAX-based methods. Best models to the right.
Fig. 15: CD Plots for dictionary-based methods. Best models to the right.
Fig. 16: CD Plots for interpretable methods. Best models to the right.
Fig. 17: CD Plots for all benchmarked methods. Best models to the right.
Fig. 18: Comparison of median F1 performance against median runtime (seconds). Best models are on the top left.
Fig. 19: Runtime boxplots (seconds), Lower is better. BORF median and interquartile range are highlighted in red.
BORF
DR CIF
HYDRA
IT
MINI ROCKET
MR HYDRA
MR SEQL
MR SQM
RDST
ROCKET
RS TSF
ST
W+ MUSE
ACSF1
0.610
0.880
0.850
0.920
0.910
0.880
0.710
0.770
0.920
0.880
0.890
0.890
0.920
Adiac
0.652
0.829
0.818
0.813
0.816
0.844
0.739
0.509
0.742
0.785
0.803
0.801
0.824
AGW-X
0.656
0.647
0.686
0.789
0.703
0.736
0.601
0.590
0.693
0.699
0.643
0.590
0.487
AGW-Y
0.706
0.661
0.750
0.817
0.769
0.791
0.680
0.624
0.754
0.774
0.711
0.626
0.563
AGW-Z
0.633
0.661
0.671
0.784
0.679
0.724
0.579
0.546
0.653
0.716
0.674
0.589
0.516
ArrowHead
0.829
0.823
0.834
0.834
0.863
0.863
0.754
0.709
0.857
0.823
0.749
0.806
0.874
Appendix
TABLE VI: Accuracy of the top 13 best-performing models on all datasets. Missing values are due to exceeded runtime limits or out-of-memory errors.
Irregular time series, characterized by non-uniform sampling intervals, missing observations, and variable lengths, are ubiquitous in healthcare, mobility, and environmental monitoring, yet effective and interpretable classifiers for this setting are limited. Existing approaches often rely on imputation, which can obscure the temporal structure of the data, or require complex neural architectures that are opaque and difficult to explain. In this work, we extend the Bag-Of-Receptive-Fields (BORF), a fast, deterministic, and interpretable transform for time series, to the irregular setting. Our key contribution is a time-weighted normalization scheme in which each observation is weighted proportionally to its associated time delta, making pattern extraction sensitive to the actual temporal distribution of samples rather than only their index position. This requires deriving an efficient sliding-window recurrence for the time-weighted standard deviation, preserving the linear time complexity of BORF. We benchmark the resulting method against state-of-the-art irregular time series classifiers on datasets from the PYRREGULAR repository, demonstrating competitive classification performance with the added benefit of human-interpretable explanations.
Francesco Spinnato
Department of Computer Science, University of Pisa, Italy · ISTI-CNR, Pisa, Italy
Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent complexity of time-series classification (TSC) and to make model decision-making processes more transparent. However, existing methods primarily focus on population-level shapelets optimized across the entire dataset, which leads to two fundamental limitations: (i) population-level patterns often misalign with instance-specific features, resulting in suboptimal performance and potentially misleading interpretations, and (ii) most methods treat shapelets as independent entities, overlooking important temporal dependencies and interactions among multiple patterns. To address these limitations, we propose INSHAPE, an interpretable TSC framework that discovers variable-length, discriminative temporal patterns specific to each time series. INSHAPE identifies these patterns as non-overlapping segments and models their temporal dependencies, thereby providing clear instance-level interpretations while achieving strong predictive performance. Furthermore, INSHAPE bridges local and global interpretability through a bottom-up approach, aggregating instance-level shapelets into prototypical (population-level) shapelets. Extensive experiments on 128 UCR and 30 UEA benchmark datasets show that INSHAPE consistently outperforms state-of-the-art shapelet-based methods while providing more intuitive and interpretable insights.
Seongjun Lee, Seokhyun Lee, Changhee Lee
Department of Artificial Intelligence, Korea University
Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology that repurposes XAI attribution methods for effective data reduction in Time Series Classification (TSC). The core challenge in modern TSC is scalability; state-of-the-art models, such as Transformers, exhibit quadratic complexity relative to sequence length and linear complexity relative to the number of channels. This renders them computationally prohibitive for massive datasets. drXAI addresses this by using a fast, GPU-accelerated classifier (Hydra) to generate local attributions. We aggregate these into global feature importance scores and employ an automated elbow-cut heuristic to select the most salient features without requiring manual thresholds. We evaluate our approach on both synthetic and real-world univariate and multivariate datasets. On synthetic benchmarks, drXAI successfully recovers ground-truth features where traditional baselines fail. On real-world data, drXAI achieves between 80% and 90% data reduction while maintaining classification accuracy comparable to models trained on the full dataset. Most importantly, we show that drXAI allows resource-intensive models like ConvTran to scale to datasets that were previously inaccessible due to memory constraints. Our results show the benefits of using XAI not just for interpretability, but as a robust tool for feature selection and scalability in time series analysis. All our code and data are openly available.
Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
School of Computer Science, University College Dublin, Ireland