Organizations: The University of Tokyo, Tokyo, Japan · Stony Brook University, SUNY Korea, Incheon, South Korea · Fuyao University of Science and Technology, Fujian, China · Renmin University of China, Beijing, China · Tsinghua University, Beijing, China · Anhui Science and Technology University, Anhui, China
Classifiers for time-series classification are commonly selected on validation data, but the temporal pattern of missing observations at deployment may differ from the pattern seen during validation. We examine whether such a mismatch affects validation-based classifier selection. In a controlled 2×2 design, validation and test sets of 64 univariate UCR datasets were masked with either random point missingness or circular block missingness at six rates from 5% to 30%, imputed by linear interpolation, and used to select among three prespecified candidates: 1NN-DTW, MiniRocket with a Ridge classifier, and a statistical-feature Random Forest. Training data remained complete, and selections made under matched and mismatched validation patterns were compared on the same masked test sets. Mismatched validation reduced the test balanced accuracy of the selected classifier by 1.14 percentage points on average (95% CI 0.79 to 1.51), with losses on 49 of the 64 datasets. The loss was negligible at 5% missingness and increased to 2.46 percentage points at 30%. It was concentrated in point-masked deployment (1.84 percentage points), where block-masked validation shifted selection away from the usually best candidate, while the effect for block-masked deployment was small and not significant. Mismatch changed the selected classifier in 35.5% of paired comparisons, but a changed selection did not always reduce performance. A supplementary analysis with non-wrapping linear blocks reproduced these findings with a larger effect (1.67 percentage points). Matching the missingness pattern of validation data to the expected deployment pattern is therefore a simple safeguard for method selection, particularly at higher missingness rates.
Figures & tables
Figure 1 : Experimental design. Candidates are fitted on the complete (unmasked) fitting subset. Validation and test sets are masked with point ( P ) or circular block ( B ) missingness and imputed by linear interpolation; the candidate with the highest validation balanced accuracy (BA) is then evaluated on the test set. Grey cells are matched and white cells mismatched conditions. ΔP and ΔB are the target-paired effects defined in Section 3.12.2 .
Δ BA (pp)
Δ SE (pp)
Comparison
Mean
95% CI
W/B/E
p
Mean
95% CI
p
Both targets
1.14
[0.79, 1.51]
49/7/8
1.8×10−9
6.4
[3.6, 9.4]
1.7×10−5
Point test ( BP vs PP )
1.84
[1.23, 2.50]
43/13/8
7.4×10−7
11.5
[7.1, 16.2]
2.5×10−5
Block test ( PB vs BB )
0.44
[ −0.11 , 1.03]
34/22/8
0.20
1.4
[ −2.3 , 5.1]
0.46
Both targets, ties excluded
1.23
[0.75, 1.78]
44/6/12
5.3×10−8
8.5
[4.5, 13.3]
6.5×10−5
Table 1 : Target-paired effect of mismatched validation in the primary experiment. Δ BA is the loss in the selected candidate’s test balanced accuracy (equivalently, the increase in regret); Δ SE is the increase in the selection-error rate. Means over 64 datasets with 95% bootstrap intervals; W/B/E counts datasets on which mismatch was worse, better, or equal; p from two-sided Wilcoxon signed-rank tests. The tie-excluded row omits target pairs with a validation tie on either side (62 datasets remain).
Δ BA (pp)
Δ SE (pp)
Rate
Mean
95% CI
W/B/E
pHolm
Mean
pHolm
Changed (%)
5%
−0.03
[ −0.09 , 0.03]
18/17/29
0.79
0.8
0.27
21.3
10%
0.61
[0.29, 0.99]
30/12/22
3.0×10−3
5.5
0.018
27.8
15%
0.91
[0.46, 1.41]
34/12/18
2.0×10−4
5.9
0.037
35.3
20%
1.19
[0.69, 1.74]
40/10/14
2.8×10−5
4.7
0.26
36.9
25%
1.69
[1.03, 2.41]
40/11/13
9.5×10−6
9.8
1.8×10−3
43.8
Table 2 : Target-paired effect by nominal missingness rate in the primary experiment, averaged over seeds and both target patterns. pHolm values are Holm-adjusted across the six rates within each metric. Changed is the percentage of target pairs in which mismatched validation selected a different candidate.
Figure 2 : Target-paired cost of mismatched validation by missingness rate in the primary experiment. (A) Loss in the selected candidate’s test balanced accuracy, equal to the increase in regret. (B) Increase in the selection-error rate. Lines show means over 64 datasets for point-masked tests ( BP vs PP ), circular-block-masked tests ( PB vs BB ), and both targets; bands are 95% bootstrap intervals over datasets. Positive values indicate that mismatch is worse.
Regret
Selection error
Selected BA
Oracle BA
Validation ties
Condition
(pp)
(%)
(%)
(%)
(%)
PP (matched)
2.06 [1.49, 2.71]
26.9 [20.5, 33.5]
79.4
81.5
20.1
BP (mismatched)
3.89 [3.03, 4.89]
38.4 [31.7, 45.2]
77.6
81.5
13.4
BB (matched)
2.49 [2.01, 2.99]
32.9 [27.7, 38.3]
67.9
70.4
13.4
PB (mismatched)
2.93 [2.27, 3.60]
34.3 [28.2, 40.4]
67.4
70.4
20.1
Table 3 : Dataset-level outcomes of the four primary conditions (validation pattern → test pattern), averaged over seeds and rates; 95% bootstrap intervals in brackets. The oracle depends only on the test pattern and the tie rate only on the validation pattern.
Figure 3 : Outcomes of the four primary conditions by missingness rate. (A) Regret against the candidate-set test oracle. (B) Selection-error rate. Solid lines are matched and dashed lines mismatched conditions; bands are 95% bootstrap intervals over 64 datasets.
Figure 4 : Candidate shares in the primary experiment, weighting datasets equally. Left: candidate selected by validation balanced accuracy. Right: best candidate on the test set (oracle); ties between k candidates give each a share of 1/k . Selection depends only on the validation pattern and the oracle only on the test pattern.
Figure 5 : Dataset-level target-paired cost of mismatched validation in the primary experiment: loss in the selected candidate’s test balanced accuracy, averaged over seeds, rates, and both target patterns. Orange: mismatch worse; blue: mismatch better; grey: no difference within 10−9 .
Outcome
Circular block
Linear block
Δ BA, both targets (pp)
1.14 [0.79, 1.51]
1.67 [1.23, 2.14]
W/B/E datasets; p
49/7/8; 1.8×10−9
54/4/6; 1.2×10−10
Δ BA, point test (pp)
1.84 [1.23, 2.50]
2.24 [1.62, 2.90]
Δ BA, block test (pp)
0.44 [ −0.11 , 1.03]
1.10 [0.45, 1.81]
p (block test)
0.20
0.013
Δ SE, both targets (pp)
6.4 [3.6, 9.4]
11.2 [7.3, 15.4]
Table 4 : Target-paired effect of mismatched validation under circular blocks (primary experiment) and non-wrapping linear blocks (supplementary experiment). Each analysis uses 64 datasets and is summarized separately; the PP condition is shared. Values are means with 95% bootstrap intervals; the trend is the mean dataset-level slope in pp of balanced accuracy per percentage point of missingness.
Figure 6 : Target-paired cost of mismatched validation under circular blocks (primary experiment) and linear blocks (supplementary experiment) by missingness rate: (A) point-masked tests, (B) block-masked tests, (C) both targets. Each analysis is summarized separately over 64 datasets with 95% bootstrap intervals; the PP condition is shared, and the analyses are not pooled.
Handling missing data in time series classification remains a significant challenge in various domains. Traditional methods often rely on imputation, which may introduce bias or fail to capture the underlying temporal dynamics. In this paper, we propose TANDEM (Temporal Attention-guided Neural Differential Equations for Missingness), an attention-guided neural differential equation framework that effectively classifies time series data with missing values. Our approach integrates raw observation, interpolated control path, and continuous latent dynamics through a novel attention mechanism, allowing the model to focus on the most informative aspects of the data. We evaluate TANDEM on 30 benchmark datasets and a real-world medical dataset, demonstrating its superiority over existing state-of-the-art methods. Our framework not only improves classification accuracy but also provides insights into the handling of missing data, making it a valuable tool in practice.
YongKyung Oh, Dong-Young Lim, Sungil Kim +1
University of California, Los Angeles, Los Angeles, CA, USA · Ulsan National Institute of Science and Technology, Ulsan, Republic of Korea
Naturally irregular time series combine asynchronous observations, missing values, unequal lengths, and nonuniform sampling, while dense adapters can discard temporal structure. We propose a mask-aware state space classifier for irregular time series (MASCIT), which supplies observation masks to the encoder and excludes invalid steps from gated temporal aggregation. Across 34 irregular time series datasets, MASCIT yielded the strongest aggregate point estimate and was the only evaluated neural model with three-seed results on every dataset. MASCIT retained the lowest point rank across six overlapping irregularity indicators, while factorial ablations favored partial over full selectivity. These results support selective state space models as effective, executable backbones for naturally irregular time series classification.
Yoo-Min Jung, Hyeon-Gi Kim, Jonghun Park
Department of Industrial Engineering Seoul National University, Seoul, Republic of Korea
Multichannel time-series classification commonly assumes synchronized sensor streams, although latency, clock drift, and preprocessing can introduce relative delays during data collection or after deployment. Existing synchronization solutions are often hardware-specific and difficult to apply retrospectively. Consequently, synchronization problems may remain undetected while classification performance is suboptimal. We introduce a classifier- and label-free diagnostic based on minimum description length (MDL). Our method applies candidate temporal shifts to sensor groups and measures how efficiently one group can be encoded through a representation of the remaining channels. An increased codelength indicates that the shift destroys shared temporal structure, whereas the minimum identifies the alignment most strongly supported by the data. Unlike learned synchronization methods, the diagnostic requires neither retraining nor a trusted aligned reference and can therefore test both training and deployment data for misalignments. Experiments on two controlled synthetic tasks and nine real-world datasets show that the metric exposes alignment structure and can recover accuracy under induced deployment drift. A whole-dataset audit further identifies stable nonzero MDL optima in established benchmarks including FordChallenge, Opportunity, PAMAP2, and UCIActivity, revealing potential systematic offsets that conventional model evaluation does not expose. Our method thus provides a general-purpose tool for detecting, understanding, and correcting temporal misalignment throughout the time-series learning pipeline. Our code is available under https://github.com/sbuschjaeger/mdl-temporal-misalignment.
Sebastian Buschjäger, Michael Frichert, Daniel Kuhse +1
Lamarr Institute, TU Dortmund University, Germany · Cyber-Physical Systems, RWTH Aachen, Germany