Multichannel time-series classification commonly assumes synchronized sensor streams, although latency, clock drift, and preprocessing can introduce relative delays during data collection or after deployment. Existing synchronization solutions are often hardware-specific and difficult to apply retrospectively. Consequently, synchronization problems may remain undetected while classification performance is suboptimal. We introduce a classifier- and label-free diagnostic based on minimum description length (MDL). Our method applies candidate temporal shifts to sensor groups and measures how efficiently one group can be encoded through a representation of the remaining channels. An increased codelength indicates that the shift destroys shared temporal structure, whereas the minimum identifies the alignment most strongly supported by the data. Unlike learned synchronization methods, the diagnostic requires neither retraining nor a trusted aligned reference and can therefore test both training and deployment data for misalignments. Experiments on two controlled synthetic tasks and nine real-world datasets show that the metric exposes alignment structure and can recover accuracy under induced deployment drift. A whole-dataset audit further identifies stable nonzero MDL optima in established benchmarks including FordChallenge, Opportunity, PAMAP2, and UCIActivity, revealing potential systematic offsets that conventional model evaluation does not expose. Our method thus provides a general-purpose tool for detecting, understanding, and correcting temporal misalignment throughout the time-series learning pipeline. Our code is available under https://github.com/sbuschjaeger/mdl-temporal-misalignment.
Figures & tables
Dataset
n
T
C
∣G∣
∣Y∣
CrowdSourced
12,289
256
14
14
2
DreamerA
170,246
256
14
14
2
DreamerV
170,246
256
14
14
2
STEW
28,512
256
14
14
2
Opportunity
17,386
100
113
8
5
PAMAP2
38,856
100
52
4
12
Table 1: Datasets used in the experiments. The table reports the number of windows n , window length T , channels C , channel groups ∣G∣ , and classes ∣Y∣ . Real-world datasets (top group) are part of MONSTER ( Dempster et al. 2025 ) .
Figure 1: Results on the synthetic experiments. The left column shows results for the sensitive datasets, and the right column shows results for the robust one. The first row depicts an example of both datasets, where class labels are depicted as shaded regions. The second row shows the alignment costs according to our metric, and the last row shows the accuracy for all four classifiers with the shaded region depicting the standard deviation over 5 cross-validation folds.
Dataset
Group
Global τG⋆
Training-fold τG⋆
QG(0) [kbit]
FordChallenge
env
6
4–8
248.5
Opportunity
r shoe
15
8–15
32.2
PAMAP2
ankle
20
8–20
807.4
UCIActivity
gyro
56
56
389.4
Skoda
s05
0
0–0
0.0
CrowdSourced
O2
0
0–0
0.0
Table 2: Sensor groups with the largest misalignment in each dataset. ∗ Reached the maximum lag tested, so a larger lag misalignment might be present in the data.
Figure 2: Accuracy of the four classifiers on 9 real-world datasets with imposed lags with and without correction. The solid lines show the average accuracy of the model across the five folds over imposed lags on the test data. The dashed line shows the accuracy after correcting the imposed lag via the alignment loss. For visual clarity, uncertainty bands are not displayed. See appendix for uncertainty bands.
Dataset
Our metric [%]
SyncNet [%]
Diff [pp]
FordChallenge
93.66
94.41
-0.75
DREAMERV
98.95
100.0
-1.05
Opportunity
94.38
100.0
-5.62
UCIActivity
94.38
100.0
-5.62
PAMAP2
89.44
100.0
-10.56
CrowdSourced
55.11
78.07
-22.96
Table 3: Baseline-relative lag correction using the label-free MDL diagnostic and learned SyncNet baseline. Columns report mean absolute residual error normalized by maximum window length without lag. Higher is better.
Setting
Value
Random seed
42+ fold index
Validation split
stratified 15% of training fold
Input normalization
per-channel training mean/std.
Training / evaluation batch
256 / 2048
Maximum epochs / early stopping
100 / patience 10
Classifier optimizer
Adam, 10−4
Table 4: Training and evaluation settings used for every dataset and fold.
Table 5: Configured non-negative lag grids and the sensor group used for correction and the SyncNet comparison. The notation a:s:b denotes values from a to b in steps of s .
Figure 3: Test accuracy under imposed sensor delays before and after label-free MDL correction. Lines show fold means and bands show one standard deviation over five folds. Solid lines are evaluated at the imposed lag. Dashed lines use the residual lag selected by the MDL alignment cost.
Figure 4: Label-free alignment cost over all evaluated real-world datasets. Lines show fold means and bands show one standard deviation. Each panel uses the classifier-selected group from Table 5 . Raw bit scales are dataset-specific.
Dataset
MDL
SyncNet
Cases
FordChallenge
2.03
1.79
33
Opportunity
4.50
0.00
22
PAMAP2
8.44
0.00
9
UCIActivity
3.60
0.00
10
Skoda
0.00
0.00
45
CrowdSourced
79.00
38.60
5
Table 6: Baseline-relative lag errors in samples (lower is better). Cases counts the jointly reachable fold–lag pairs.
Dataset
Nominal
Lagged
MDL
SyncNet
FordChallenge
86.5
86.5
86.5
86.5
Opportunity
81.6
81.1
81.2
81.6
PAMAP2
74.0
72.6
72.8
74.0
UCIActivity
92.4
89.9
89.9
92.4
Skoda
91.4
91.1
91.4
91.4
CrowdSourced
67.1
66.0
67.3
66.3
Table 7: Classifier accuracy after lag correction [%]. Values average CNN, ResNet, late fusion, and Magnitude CNN over five folds and the jointly reachable imposed lags.
Dataset
Group
U lag
C lag
FordChallenge
vehicle
0 [0–8]
6 [0–8]
Opportunity
l shoe
5 [5–15]
15 [0–15]
PAMAP2
ankle
20 [8–20]
3 [0–5]
UCIActivity
gyro
56 [56–56]
64 [64–64]
Skoda
s06
0 [0–0]
0 [0–0]
CrowdSourced
F7
80 [80–80]
65 [65–80]
Table 8: Preferred lags for the label-free (U) and class-conditional (C) metrics, reported as median [minimum–maximum] over five folds.
Dataset
U AG
C AG
FordChallenge
15.9
17.2
Opportunity
94.1
122.2
PAMAP2
278.1
51.7
UCIActivity
167.5
166.1
Skoda
28.0
82.8
CrowdSourced
1.7
25.2
Table 9: Mean raw alignment cost AG [kbit] for the label-free (U) and class-conditional (C) metrics.
Classifiers for time-series classification are commonly selected on validation data, but the temporal pattern of missing observations at deployment may differ from the pattern seen during validation. We examine whether such a mismatch affects validation-based classifier selection. In a controlled 2×2 design, validation and test sets of 64 univariate UCR datasets were masked with either random point missingness or circular block missingness at six rates from 5% to 30%, imputed by linear interpolation, and used to select among three prespecified candidates: 1NN-DTW, MiniRocket with a Ridge classifier, and a statistical-feature Random Forest. Training data remained complete, and selections made under matched and mismatched validation patterns were compared on the same masked test sets. Mismatched validation reduced the test balanced accuracy of the selected classifier by 1.14 percentage points on average (95% CI 0.79 to 1.51), with losses on 49 of the 64 datasets. The loss was negligible at 5% missingness and increased to 2.46 percentage points at 30%. It was concentrated in point-masked deployment (1.84 percentage points), where block-masked validation shifted selection away from the usually best candidate, while the effect for block-masked deployment was small and not significant. Mismatch changed the selected classifier in 35.5% of paired comparisons, but a changed selection did not always reduce performance. A supplementary analysis with non-wrapping linear blocks reproduced these findings with a larger effect (1.67 percentage points). Matching the missingness pattern of validation data to the expected deployment pattern is therefore a simple safeguard for method selection, particularly at higher missingness rates.
Ruiqi Zhao, Zishun Yuan, Zhentao Wang +3
The University of Tokyo, Tokyo, Japan · Stony Brook University, SUNY Korea, Incheon, South Korea · Fuyao University of Science and Technology, Fujian, China +3
Reconstruction errors in multivariate time-series anomaly detection may not reliably distinguish abnormal behavior from benign deviations. Language-derived semantics offer complementary context, but existing multimodal approaches may rely on time-associated paired textual information that is difficult to obtain consistently and is not provided by standard multivariate time-series anomaly detection benchmarks. This setting poses two challenges: (1) conditioning masked reconstruction on window-specific semantics without exposing exact numerical targets or anomaly-specific cues, and (2) using a window-independent concept of normality as a complementary semantic reference rather than an independent anomaly detector. We propose LLM-Enhanced Alignment and Reconstruction with Normality Guidance for Time Series (LEARN-TS), which uses a frozen language model to construct two role-separated semantic representations without requiring temporally paired external text. Window-specific observation semantics encode temporal and cross-variable context without exact numerical values to guide channel-shared patch-masked reconstruction. A fixed, dataset-agnostic normality prompt provides a window-independent semantic reference for aligning normal representations and estimating normality discrepancy. At inference, masking each temporal patch once yields timestamp-level reconstruction evidence, conditionally modulated by discrepancy from a separate unmasked view. Across four benchmarks, LEARN-TS achieves the highest mean performance in 13 of 16 dataset-metric comparisons. Controlled ablations examine observation conditioning, joint normality alignment and scoring, and reference content, showing dataset-dependent ranking benefits and modest average gains from semantic over random references.
Jahyeob Koo, Kio Yun, Byoungmo Koo +1
Department of Industrial and Management Engineering, Korea University Seoul, Republic of Korea
Temporal classification errors are often treated as representation failures, but they can also arise from how available evidence is converted into decisions. This paper proposes a representation--calibration decomposition for temporal classification. We keep a trained native classifier frozen and separate two inference-time interventions: a conservative residual multi-scale branch that adds auxiliary logits to the native prediction, and a post-hoc branch-aware calibrator that recombines native and residual evidence at decision time. This design distinguishes missing temporal evidence from underused decision-level evidence without retraining the backbone. Across FI-2010, PTB-XL, UCI-HAR, MHEALTH, and HARTH, we find that gains are strongly regime-dependent. Residual multi-scale evidence is most useful in noisy or representation-limited settings, especially short-horizon FI-2010 and weaker recurrent backbones, while branch-aware calibration helps when native and auxiliary logits contain complementary evidence not fully exploited by the raw decision rule. Near-saturated settings show limited gains from either intervention. These results suggest that temporal classification should be understood not only as representation learning, but also as the problem of trusting, combining, and calibrating evidence from multiple views.
Arthur Chagas, Arthur Buzelin, Yan Aquino +4
Department of Computer Science (DCC), Universidade Federal de Minas Gerais (UFMG)