Organizations: The University of Tokyo, Tokyo, Japan · Stony Brook University, SUNY Korea, Incheon, South Korea · Fuyao University of Science and Technology, Fujian, China · Renmin University of China, Beijing, China · Tsinghua University, Beijing, China · Anhui Science and Technology University, Anhui, China
Classifiers for time-series classification are commonly selected on validation data, but the temporal pattern of missing observations at deployment may differ from the pattern seen during validation. We examine whether such a mismatch affects validation-based classifier selection. In a controlled 2×2 design, validation and test sets of 64 univariate UCR datasets were masked with either random point missingness or circular block missingness at six rates from 5% to 30%, imputed by linear interpolation, and used to select among three prespecified candidates: 1NN-DTW, MiniRocket with a Ridge classifier, and a statistical-feature Random Forest. Training data remained complete, and selections made under matched and mismatched validation patterns were compared on the same masked test sets. Mismatched validation reduced the test balanced accuracy of the selected classifier by 1.14 percentage points on average (95% CI 0.79 to 1.51), with losses on 49 of the 64 datasets. The loss was negligible at 5% missingness and increased to 2.46 percentage points at 30%. It was concentrated in point-masked deployment (1.84 percentage points), where block-masked validation shifted selection away from the usually best candidate, while the effect for block-masked deployment was small and not significant. Mismatch changed the selected classifier in 35.5% of paired comparisons, but a changed selection did not always reduce performance. A supplementary analysis with non-wrapping linear blocks reproduced these findings with a larger effect (1.67 percentage points). Matching the missingness pattern of validation data to the expected deployment pattern is therefore a simple safeguard for method selection, particularly at higher missingness rates.
Figures & tables
Figure 1 : Experimental design. Candidates are fitted on the complete (unmasked) fitting subset. Validation and test sets are masked with point ( P ) or circular block ( B ) missingness and imputed by linear interpolation; the candidate with the highest validation balanced accuracy (BA) is then evaluated on the test set. Grey cells are matched and white cells mismatched conditions. ΔP and ΔB are the target-paired effects defined in Section 3.12.2 .
Δ BA (pp)
Δ SE (pp)
Comparison
Mean
95% CI
W/B/E
p
Mean
95% CI
p
Both targets
1.14
[0.79, 1.51]
49/7/8
1.8×10−9
6.4
[3.6, 9.4]
1.7×10−5
Point test ( BP vs PP )
1.84
[1.23, 2.50]
43/13/8
7.4×10−7
11.5
[7.1, 16.2]
2.5×10−5
Block test ( PB vs BB )
0.44
[ −0.11 , 1.03]
34/22/8
0.20
1.4
[ −2.3 , 5.1]
0.46
Both targets, ties excluded
1.23
[0.75, 1.78]
44/6/12
5.3×10−8
8.5
[4.5, 13.3]
6.5×10−5
Table 1 : Target-paired effect of mismatched validation in the primary experiment. Δ BA is the loss in the selected candidate’s test balanced accuracy (equivalently, the increase in regret); Δ SE is the increase in the selection-error rate. Means over 64 datasets with 95% bootstrap intervals; W/B/E counts datasets on which mismatch was worse, better, or equal; p from two-sided Wilcoxon signed-rank tests. The tie-excluded row omits target pairs with a validation tie on either side (62 datasets remain).
Δ BA (pp)
Δ SE (pp)
Rate
Mean
95% CI
W/B/E
pHolm
Mean
pHolm
Changed (%)
5%
−0.03
[ −0.09 , 0.03]
18/17/29
0.79
0.8
0.27
21.3
10%
0.61
[0.29, 0.99]
30/12/22
3.0×10−3
5.5
0.018
27.8
15%
0.91
[0.46, 1.41]
34/12/18
2.0×10−4
5.9
0.037
35.3
20%
1.19
[0.69, 1.74]
40/10/14
2.8×10−5
4.7
0.26
36.9
25%
1.69
[1.03, 2.41]
40/11/13
9.5×10−6
9.8
1.8×10−3
43.8
Table 2 : Target-paired effect by nominal missingness rate in the primary experiment, averaged over seeds and both target patterns. pHolm values are Holm-adjusted across the six rates within each metric. Changed is the percentage of target pairs in which mismatched validation selected a different candidate.
Figure 2 : Target-paired cost of mismatched validation by missingness rate in the primary experiment. (A) Loss in the selected candidate’s test balanced accuracy, equal to the increase in regret. (B) Increase in the selection-error rate. Lines show means over 64 datasets for point-masked tests ( BP vs PP ), circular-block-masked tests ( PB vs BB ), and both targets; bands are 95% bootstrap intervals over datasets. Positive values indicate that mismatch is worse.
Regret
Selection error
Selected BA
Oracle BA
Validation ties
Condition
(pp)
(%)
(%)
(%)
(%)
PP (matched)
2.06 [1.49, 2.71]
26.9 [20.5, 33.5]
79.4
81.5
20.1
BP (mismatched)
3.89 [3.03, 4.89]
38.4 [31.7, 45.2]
77.6
81.5
13.4
BB (matched)
2.49 [2.01, 2.99]
32.9 [27.7, 38.3]
67.9
70.4
13.4
PB (mismatched)
2.93 [2.27, 3.60]
34.3 [28.2, 40.4]
67.4
70.4
20.1
Table 3 : Dataset-level outcomes of the four primary conditions (validation pattern → test pattern), averaged over seeds and rates; 95% bootstrap intervals in brackets. The oracle depends only on the test pattern and the tie rate only on the validation pattern.
Figure 3 : Outcomes of the four primary conditions by missingness rate. (A) Regret against the candidate-set test oracle. (B) Selection-error rate. Solid lines are matched and dashed lines mismatched conditions; bands are 95% bootstrap intervals over 64 datasets.
Figure 4 : Candidate shares in the primary experiment, weighting datasets equally. Left: candidate selected by validation balanced accuracy. Right: best candidate on the test set (oracle); ties between k candidates give each a share of 1/k . Selection depends only on the validation pattern and the oracle only on the test pattern.
Figure 5 : Dataset-level target-paired cost of mismatched validation in the primary experiment: loss in the selected candidate’s test balanced accuracy, averaged over seeds, rates, and both target patterns. Orange: mismatch worse; blue: mismatch better; grey: no difference within 10−9 .
Outcome
Circular block
Linear block
Δ BA, both targets (pp)
1.14 [0.79, 1.51]
1.67 [1.23, 2.14]
W/B/E datasets; p
49/7/8; 1.8×10−9
54/4/6; 1.2×10−10
Δ BA, point test (pp)
1.84 [1.23, 2.50]
2.24 [1.62, 2.90]
Δ BA, block test (pp)
0.44 [ −0.11 , 1.03]
1.10 [0.45, 1.81]
p (block test)
0.20
0.013
Δ SE, both targets (pp)
6.4 [3.6, 9.4]
11.2 [7.3, 15.4]
Table 4 : Target-paired effect of mismatched validation under circular blocks (primary experiment) and non-wrapping linear blocks (supplementary experiment). Each analysis uses 64 datasets and is summarized separately; the PP condition is shared. Values are means with 95% bootstrap intervals; the trend is the mean dataset-level slope in pp of balanced accuracy per percentage point of missingness.
Figure 6 : Target-paired cost of mismatched validation under circular blocks (primary experiment) and linear blocks (supplementary experiment) by missingness rate: (A) point-masked tests, (B) block-masked tests, (C) both targets. Each analysis is summarized separately over 64 datasets with 95% bootstrap intervals; the PP condition is shared, and the analyses are not pooled.