Organizations: Defense Innovation Institute, Academy of Military Science, Beijing, China · School of Software Engineering, South China University of Technology, Guangzhou, China
Wearable human-activity recognition (HAR) models operate across sensors, subjects, and backbones, yet a smooth waveform may appear temporal while exploiting a persistent sensor offset primarily. We introduce SpectrumAudit, a label-sealed audit that fits a phase-randomized full-window stimulus on calibration windows from subjects held out from training and testing. After selection, it replays its exact DC projection and budget-constrained zero-mean residual on the same frozen victim without refitting. Across 27 victims from three datasets and three backbones, the selected waveforms cause 2.87-40.83-point three-phase robust accuracy losses. Under this replay budget, DC is more damaging than AC on 24/27 victims and recovers at least 90% of the full drop on 22/27; all 5 failures occur on WISDM. In a held-out UTD-MHAD check, the selected waveform causes 13.49-pp accuracy and 11.68-pp macro-F1 losses, versus -0.66 pp for matched random changes. The audit diagnoses offset versus zero-mean variation under a common peak-budget cap. The code will be released upon acceptance.
Figures & tables
Method
Access
Drop (pp)
Random universal
none
-0.08 ± 0.29
Pseudo-label UAP–CE [ 16 ]
target
23.94 ± 9.16
Label-aware smooth UAP [ 14 ]
target+Y
15.99 ± 8.45
Auxiliary-ensemble transfer [ 9 ]
source
21.46 ± 9.93
SpectrumAudit (ours)
target
23.98 ± 8.10
Table 1: Matched controls on nine development settings (one fixed victim per dataset–backbone cell; mean ± SD across cells). Literature-derived rows are protocol-matched adaptations, not reproductions of published rates. “target” means the frozen target checkpoint is used during fitting; “source” means only the two non-target backbones are used; “target+Y” uses calibration labels in fitting and selection. Drops are in percentage points; Y denotes calibration labels.
Dataset
Victim
Clean (%)
Full-window
Const. ctrl.
AC ctrl.
ΔF−C
C > AC
PAMAP2
FCN
60.7
19.9 ± 11.7
21.8 ± 9.2
1.1 ± 2.0
-1.9 ± 2.8
3/3
PAMAP2
ResCNN
58.0
26.6 ± 4.5
26.1 ± 4.7
2.5 ± 2.4
0.6 ± 1.4
3/3
PAMAP2
Tr.
64.3
21.7 ± 8.2
23.2 ± 9.1
4.7 ± 4.1
-1.5 ± 1.1
3/3
WISDM
FCN
91.6
16.4 ± 13.3
14.2 ± 14.2
12.9 ± 7.8
2.1 ± 3.8
1/3
WISDM
ResCNN
91.0
23.2 ± 16.2
21.1 ± 17.4
14.8 ± 0.4
2.1 ± 3.9
1/3
WISDM
Tr.
85.3
12.1 ± 5.3
15.0 ± 9.3
9.1 ± 5.1
-3.0 ± 13.4
1/3
Table 2: Mainline and independently optimized controls on 27 frozen victims ( ϵ=0.30 ); full-window, Const., AC, and ΔF−C entries are mean ± sample SD over three seeds, whereas Clean is the three-seed mean; accuracy drops are in percentage points. Const. and AC are independently optimized mechanism controls, not matched attack rankings or post-selection projections. ΔF−C is full minus constant in three-phase worst-case drop; C > AC counts constant wins over AC.
Audio carries rich cues about human activities, and microphones are already built into most wearable devices. However, microphones also capture speech, and this privacy risk limits their use in Human Activity Recognition (HAR). We present AnomaSense, a sensor activation approach for wrist wearables that keeps the microphone off by default and turns it on for at most one second when an unsupervised anomaly detector flags an IMU segment that is likely to produce sound. The captured audio is further masked before it reaches the recognition model. We study 20 activities from 15 participants, organized into five groups in which activities share similar wrist motion but differ in the object or material involved. With IMU data alone, our recognition model reaches 78.98% accuracy in leave-one-participant-out validation. With the short, masked audio windows added, accuracy reaches 96.89% with no masking and stays above 86% when 90% of each one-second audio window is removed. On the same data, the anomaly detector triggers the microphone with 86.46% precision and 74.28% recall relative to sound events. We also report a small preliminary check of automatic speech recognition on masked speech, which shows that contiguous masking degrades recognition far more than point-wise masking at the same masking ratio. Our evaluation is a controlled, offline feasibility study. We describe the threat model, what the approach does and does not protect, and the steps needed before deployment.
Xue Wang, Yang Zhang
University of California, Los Angeles Los Angeles, CA, USA
Wearable human activity recognition (WHAR) models often suffer from performance degradation under real-world cross-user distribution shifts. Test-time adaptation (TTA) mitigates this degradation by adapting models online using unlabeled test streams, yet existing methods largely inherit assumptions from vision tasks and underexploit the inherent inter-window temporal structure in WHAR streams. In this paper, we revisit such temporal structure as a feature-conditioned inference signal rather than merely an output-space smoothing prior. We derive the insight that temporal continuity and observation-induced feature deviations provide complementary cues for determining when to preserve or release temporal inertia and where to route prediction refinement during likely transitions. Building upon this insight, we propose SIGHT, a lightweight and backpropagation-free TTA framework for WHAR, enabling real-time edge deployment. SIGHT estimates predictive surprise by comparing the current feature with a prototype-based expected state, and then uses the resulting feature deviation to guide geometry-aware transition routing based on prototype alignment and stream-level marginal habit tracking. Evaluations on real-world datasets confirm that SIGHT outperforms existing TTA baselines while reducing computational and memory costs.
Zishu Zhou, Zaipeng Xie, Xuanyao Jie
College of Computer Science and Software Engineering, Hohai University
With each sensing modality exhibiting inherent strengths and limitations, multi-modal approaches for wearable Human Activity Recognition (HAR) are becoming increasingly relevant -- particularly for recognizing Activities of Daily Living (ADLs), where individual modalities often produce ambiguous signals for similar or complex activities. This work introduces HARMES, a multi-modal wearable dataset combining three wrist-recorded modalities: motion sensing via an Inertial Measurement Unit (IMU), atmospheric environmental sensors (humidity, temperature, and pressure), and audio. Collected from 20 participants performing household activities in their own homes, HARMES totals over 80 hours of recorded data, with approximately three hours of labeled activity data per participant across 15 ADL classes. To the best of our knowledge, HARMES is the first dataset to combine this particular sensor trio, and it is nearly six times larger than the previously largest wrist-inertial-acoustic HAR dataset. In an extensive benchmark, we evaluate cross-subject generalization and conduct an ablation study revealing that modality contributions are activity-dependent and can provide complementary value, particularly for activities that are ambiguous from motion data alone. HARMES is freely available at Zenodo, alongside example code for loading the dataset and training models on GitHub.
Robin Burchard, Pascal-André Brückner, Marius Bock +2
University of Siegen, Siegen, Germany · University of Bonn and & Lamarr Institute for Machine Learning and Artificial Intelligence, Bonn, Germany