TopTimeNet: Topologically-assisted time-series classification model
Authors: Sharareh Sayyad, Sophia Bazzi
Organizations: Department of Mathematics and Statistics, Washington State University, Pullman, Washington 99164-3113, USA · European Molecular Biology Laboratory, EMBL Hamburg, c/o DESY, Building 25A, Notkestraße 85, 22603 Hamburg, Germany
Distinguishing periodic from chaotic dynamics in a time series is a fundamental challenge in both physics and engineering. Yet, end-to-end learned architectures must discover both a representation and a decision boundary from data, at substantial cost. We introduce TopTimeNet, which decouples these tasks: a fixed, non-learned stage extracts a 42-dimensional geometric and topological descriptor from Takens delay embeddings and persistent homology, and a lightweight learnable stage performs classification. On a benchmark of 49 nonlinear dynamical systems, a 1,638-parameter configuration matches the mean accuracy of one with 33× more trainable parameters. Additionally, this approach delivers mean accuracy comparable to convolutional neural networks and surpasses the average performance of converged Transformer models, while requiring three to four orders of magnitude fewer trainable parameters. Robustness also depends sharply on where noise is introduced: TopTimeNet degrades gracefully under perturbations to its precomputed features, but degrades sharply when noise is introduced into the raw signal and the full feature-extraction pipeline is recomputed, showing that robustness to perturbations of the precomputed features does not imply robustness of the complete raw-signal-to-prediction pipeline. These results show that decoupling fixed geometric and topological feature construction from a lightweight discriminative stage can achieve comparable classification accuracy with substantially fewer trainable parameters.
Figures & tables
Figure 1: Schematic illustration of the Top Time Net algorithm. The segmented time-series data first pass through a non-trainable feature-extraction stage of the algorithm, where point clouds are generated and geometric and topological statistics are extracted. This process creates a complete feature set that will be used in subsequent steps. The five groups of extracted features are then processed through a Group Projector and combined using a Fusion Layer. Finally, the combined features are classified using a Classifier Head to produce the final class logits.
Figure 2: Time-dependent θ1 and its Takens embedding. Left panels present θ1(t) for the periodic (top) and chaotic (bottom) regimes. Right panels show the associated point clouds obtained via a Takens embedding with time delay τ=20 and embedding dimension demb=2 .
Feature
Periodic
Chaotic
Diameter
1.1552
3.2792
Correlation-dimension proxy
1.0672
1.6585
Mean nearest-neighbor dist.
0.0014
0.0410
Std. nearest-neighbor dist.
0.0005
0.0293
Table 1: Geometric branch feature values for the double pendulum, periodic vs. chaotic regime.
Figure 3: Persistence diagrams for the two tracked homology dimensions. H0 (connected components, circles) and H1 (loops, triangles) features are shown for the periodic (left) and chaotic (right) point clouds of Fig. 2 ; note the different axis ranges between panels. Points farther from the diagonal correspond to more persistent topological features.
Feature
Periodic
Chaotic
Entropy, H0
5.5456
6.6991
Entropy, H1
0.0379
5.0058
Table 2: Entropy branch feature values for the double pendulum, periodic vs. chaotic regime.
H0
H1
Feature
Periodic
Chaotic
Periodic
Chaotic
Maximum lifetime
0.9532
0.4576
0.9348
0.1533
Total lifetime
4.8949
52.1417
0.9386
6.0242
Dominance ratio
0.1947
0.0088
0.9959
0.0254
Coefficient of variation
6.1556
0.6769
5.0775
1.0785
Number of significant lifetimes
1
836
1
169
Table 3: Lifetime branch feature values for the double pendulum, periodic vs. chaotic regime.
Figure 4: Diameter-normalized Betti curves β(ε) for H0 (top) and H1 (bottom), periodic (left) and chaotic (right). Note the different vertical scales between panels.
Figure 5: Persistence images for H0 (top) and H1 (bottom), periodic (left) and chaotic (right). Color indicates pixel intensity (note the different color scales between panels); axes are birth and persistence pixel bins.
H0
H1
Feature
Periodic
Chaotic
Periodic
Chaotic
Max
980.0000
980.0000
1.0000
49.0000
Mean
20.4000
25.3800
0.8000
1.6400
Standard deviation
137.0863
140.7353
0.4000
7.1882
Peak position
0.0000
0.0000
0.0204
0.0408
Turning points
0
0
0
1
Table 4: Betti curve branch feature values for the double pendulum, periodic vs. chaotic regime.
H0
H1
Feature
Periodic
Chaotic
Periodic
Chaotic
Total mass
1.2801
34.7081
0.2390
5.0567
Max pixel
0.0125
0.1561
0.0030
0.0173
Active fraction
0.2500
0.5000
0.3000
0.6000
Birth centroid
0.5000
0.5000
0.5000
0.4265
Persistence centroid
0.2542
0.3124
0.8928
0.4860
Table 5: Persistence image branch feature values for the double pendulum, periodic vs. chaotic regime.
Small model
Large model
Trainable parameters
1,638
54,886
Embedding dim. D
16
128
Fusion strategy
Bilinear
Bilinear
Classifier head
()
(64,32,16)
Activation
GELU
Leaky ReLU
Dropout
0.05
0.0
Table 6: Comparison of the small and large Top Time Net configurations selected by the hyperparameter search of Sec. II.3.2 . All test metrics are mean ± std over 30 independent training runs. Both configurations use bilinear fusion (Sec. II.2.2 ), which has no rank or attention hyperparameters; the rank/attention-layer/attention-head values sampled elsewhere in the search do not apply to either selected configuration.
Top Time Net (small)
CNN ( n=10 )
Transformer ( n=7 )
Trainable parameters
1,638
1,824,898
33,435,570
Accuracy
0.9758±0.0029
0.9708±0.0110
0.9412±0.0133
F1 score
0.9758±0.0029
0.9708±0.0110
0.9412±0.0133
G-mean
0.9757±0.0029
0.9708±0.0111
0.9411±0.0133
Precision
0.9758±0.0029
0.9708±0.0110
0.9412±0.0133
Recall
0.9758±0.0029
0.9710±0.0109
0.9416±0.0132
Table 7: Top Time Net (small configuration) versus CNN and Transformer baselines. Top Time Net and CNN report mean ± std over 30 and 10 independent training runs, respectively, with no failed or collapsed runs in either case. For the Transformer, 10 independent training attempts were performed; the reported mean ± std values are computed over the 7 runs that converged, while the remaining 3 attempts remained at a near-chance validation-accuracy plateau and are reported as non-converged (see text).
Figure 6: Noise robustness of all three models. (a) Test accuracy and (b) Expected Calibration Error (ECE, M=10 equal-width bins), both as a function of the Gaussian noise standard deviation σ . For Top Time Net, noise is injected at two distinct points: directly into the precomputed 42 -dimensional feature vector of Sec. II.2.1 (“feature-level”), and into the raw time series itself, with the entire feature-extraction pipeline, Takens embedding, persistent homology, and all five feature branches, recomputed from the corrupted signal (“raw-signal”). The CNN and Transformer baselines have no intermediate feature representation to perturb separately, so only their raw-signal curves are shown. Thick lines show the mean over independent training runs ( n=30 for Top Time Net, n=10 for the CNN, n=7 for the Transformer’s converged runs), shaded bands show ±1 standard deviation, and thin lines show individual runs.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Raw double-pendulum time series under increasing Gaussian noise. Top: periodic. Bottom: chaotic.
Figure 8: Takens-embedded point clouds under raw-signal Gaussian noise. Top: periodic. Bottom: chaotic.
Figure 10: Relative feature-vector drift ∥fσ−f0∥2/∥f0∥2 under raw-signal Gaussian noise , for the periodic and chaotic double-pendulum examples of Figs. 7 – 9 .
Scientific time series often encode predictive geometric structure, including connectivity, cycles, shell-like geometry, directional changes, and nonlinear neighborhoods, that standard dot-product attention does not explicitly represent. We introduce a topology-aware attention framework that adds such structure to attention logits using persistent homology (H0-H2), anchored Euler characteristic transforms, and kernel-Hilbert channels. A validation-gated local residual captures local topological signals, including a Zeng-style local H0 component, only when held-out validation data support the correction. Exact Vietoris-Rips computations and smooth topological surrogates are evaluated under a no-leakage protocol with train-only calibration, validation-only selection, and test-only reporting. We evaluate guarded topology-aware variants across three architecture families: lightweight attention/Ridge, PatchTSTForRegression, and TimeSeriesTransformerForPrediction. Experiments include synthetic benchmarks isolating higher-order topology and real datasets covering CO2, S&P 500 return-window geometry, and NASA IMS bearing degradation. The audit uses matched paired comparisons across seven dataset units, three random seeds, and three chronological splits, giving 63 paired units per architecture and 189 paired units overall. Topology-aware models show positive paired effects when geometry is predictive, with heterogeneous magnitude across datasets and architectures. Lightweight attention/Ridge improves in 46 of 63 units, with mean relative RMSE reduction of 12.5% and paired randomization p=7.2e-4; PatchTST improves in 33 units and retains the baseline in 20 units, with 23.5% reduction and p=3.5e-5; and TimeSeriesTransformer improves in 47 units, with 47.8% reduction and p<1e-4. The results support topology as a validation-selected, architecture-compatible inductive bias.
Usef Faghihi, Amir Saki
Department of Mathematics and computer science · University of Quebec Trois Rivieres
Deep learning-based models have achieved state-of-the-art performance in Time Series Forecasting (TSF), yet their evaluation remains dominated by pointwise error metrics such as Mean Squared Error (MSE), which quantify numerical accuracy but overlook structural properties of the forecast signal, including recurrent dynamics, oscillatory behavior, and phase alignment. As a result, forecasts exhibiting over-smoothing, phase shifts, or frequency distortions may achieve favorable error scores despite substantial structural degradation. To address this limitation, we propose TopoCast, a topology-driven framework for evaluating structural fidelity in TSF. TopoCast reconstructs phase-space representations of forecast and ground-truth sequences using Takens delay embedding and applies persistent homology to characterize their intrinsic dynamics. We derive four complementary topological fidelity measures from persistence diagrams and aggregate them into a Topological Fidelity Score (TFS). We further introduce dominant cycle overlap, a novel metric that maps persistent topological features to the temporal domain to assess whether dominant oscillatory patterns occur at the correct time points. Combined with TFS, this yields the Localized Topological Fidelity Score (LTFS), a phase-aware measure that captures temporal localization errors invisible to existing evaluation metrics. Experiments on five Transformer architectures across three real-world benchmark datasets demonstrate that models with similar forecasting errors can exhibit markedly different structural fidelity profiles, revealing failure modes overlooked by conventional evaluation and highlighting the value of topology-aware forecast assessment.
Sandeepa Weerasekara, Sandareka Wickramanayake
Department of Computer Science and Engineering, University of Moratuwa, Katubedda, Moratuwa, 10400, Sri Lanka.
We study the grokking phenomenon through the lens of topology. Using persistent homology on point clouds derived from the embedding matrices of a range of models trained on modular arithmetic with varying primes, we identify a clear and consistent topological signature of grokking: a sharp increase in both the maximum and total persistence of first homology (H1). Persistence diagrams reveal the emergence of a dominant long-lived topological feature together with increasingly structured secondary features, reflecting the underlying cyclic structure of the task. Compared to existing spectral and geometric diagnostics -- specifically, Fourier analysis and local intrinsic dimension -- persistent homology provides a unified geometric and topological characterization of representation learning, capturing both local and global multi-scale structure. Ablations across data regimes and control settings show that these topological transitions are tied to generalization rather than memorization. Our results suggest that persistent homology offers a principled and interpretable framework for analyzing how neural networks internalize latent structure during training.
Yifan Tang, Qiquan Wang, Inés García-Redondo +1
Imperial College London · Queen Mary University of London · University of Fribourg