Audio carries rich cues about human activities, and microphones are already built into most wearable devices. However, microphones also capture speech, and this privacy risk limits their use in Human Activity Recognition (HAR). We present AnomaSense, a sensor activation approach for wrist wearables that keeps the microphone off by default and turns it on for at most one second when an unsupervised anomaly detector flags an IMU segment that is likely to produce sound. The captured audio is further masked before it reaches the recognition model. We study 20 activities from 15 participants, organized into five groups in which activities share similar wrist motion but differ in the object or material involved. With IMU data alone, our recognition model reaches 78.98% accuracy in leave-one-participant-out validation. With the short, masked audio windows added, accuracy reaches 96.89% with no masking and stays above 86% when 90% of each one-second audio window is removed. On the same data, the anomaly detector triggers the microphone with 86.46% precision and 74.28% recall relative to sound events. We also report a small preliminary check of automatic speech recognition on masked speech, which shows that contiguous masking degrades recognition far more than point-wise masking at the same masking ratio. Our evaluation is a controlled, offline feasibility study. We describe the threat model, what the approach does and does not protect, and the steps needed before deployment.
Figures & tables
Figure 1. Overview of the AnomaSense pipeline. Left: unsupervised activity segmentation. An LSTM variational autoencoder trained only on IMU data from non-audible motion reconstructs the incoming six channel IMU stream. The reconstruction error is smoothed and thresholded, and points above the threshold are anomaly points. Right: fine-grained activity recognition. A one second IMU window centered on the anomaly point goes to a DeepConvLSTM branch, and a one second audio window requested at the anomaly point is masked and goes to an AudioCNN branch. The two feature vectors are fused with self-attention and passed to a classifier or, for unseen activities, a clustering step. Block diagram with two shaded regions. The left region, labeled Unsupervised Activity Segmentation, shows six channel IMU signals entering an LSTM encoder, a latent vector Z, and an LSTM decoder. The reconstructed signal and the input are compared with mean squared error to give a loss curve with dots marking anomalies. The right region, labeled Fine-Grained Activity Recognition, shows highlighted IMU activity segments entering a DeepConvLSTM block, a box labeled Request Acoustic with Restriction leading to a masked audio waveform that enters an AudioCNN block, both feature vectors joined by a self-attention block, and a final scatter plot of clustered activity labels.
System
Audio rate
IMU rate
Sensor usage / label
Samples / label
Activities
Accuracy
AudioIMU ( Liang et al., 2022 )
no audio
50 Hz
10 s
500 (IMU)
23 coarse (e.g., write, chop)
74.4%
Fine-Grained Hand Activity ( Laput and Harrison, 2019 )
no audio
4,000 Hz
3 s
12,000 (IMU)
25 fine-grained (e.g., clap, wash)
95.2%
GestEar ( Becker et al., 2019 )
11,025 Hz
200 Hz
300 ms
3,308 (audio), 60 (IMU)
9 fine-grained (e.g., knock left/right)
97.2%
SAMoSA ( Mollyn et al., 2022 )
1,000 Hz
50 Hz
3 s (audio), 2 s (IMU)
3,000 (audio), 100 (IMU)
26 coarse (e.g., drill, pour, knock)
83.2% (context independent)
AnomaSense , no mask
44,100 Hz
50 Hz
1 s (audio), 1 s (IMU)
44,100 (audio), 50 (IMU)
5 groups × 4 fine-grained (e.g., hammering four materials, swinging four rackets)
96.89%
AnomaSense , contiguous 50% †
44,100 Hz
50 Hz
500 ms (audio), 1 s (IMU)
22,050 (audio), 50 (IMU)
92.33%
Table 1. Sensor usage and reported performance of related wearable IMU and acoustic HAR systems, and of AnomaSense configurations under different masking settings. Numbers for prior systems are as reported in the cited papers. Configurations marked with † use contiguous masking within the one second audio window, so the microphone is effectively on for the stated fraction of the second. Configurations marked with ∗ use point-wise masking, whose effect is comparable to sampling at the stated effective rate.
Figure 2. Two masking methods applied to a speech waveform. (A) Point-wise masking sets individual samples to zero at random. (B) Contiguous masking sets one continuous segment of the window to zero, starting at a random position. Two panels of audio waveforms. Panel A shows a speech waveform where scattered individual samples have been zeroed, leaving a waveform with the same overall shape but visible gaps at the sample level. Panel B shows the same waveform with one long continuous block set to zero, leaving speech only outside that block.
Method
Masked Data
CER
WER
No Mask
-
11.54%
15.46%
Point-wise Mask
80%
16.57%
22.70%
90%
30.64%
39.00%
95%
31.53%
40.33%
97%
39.11%
48.50%
Contiguous Mask
50%
56.62%
69.41%
Table 2. Speech recognition error under different masking methods. Recognizer trained and tested on masked speech, four speakers, leave-one-speaker-out.
Figure 3. Activity segmentation output for the 20 activities, one example recording each, grouped by row. For each activity, the top trace shows the six IMU channels with the one second IMU windows selected by the detector shaded in orange. The middle trace is the smoothed reconstruction error, with small dots marking detected anomaly points. The bottom trace is the audio waveform with the requested one second audio windows shaded in blue. The dots are small at this scale. For clearer illustration this figure uses 0.5 second windows.
Figure 4. Leave-one-participant-out recognition accuracy. (A) Trigger-selected activity segments versus randomly placed segments, for IMU only, audio only, and IMU plus audio, without masking. (B) Point-wise versus contiguous masking at 50%, 80%, and 90% mask factors, IMU plus audio. Error bars are standard deviation across held-out participants.
Mask
Factor
Accuracy
Std.
No Mask
-
96.89%
2.28%
Point-wise Mask
50%
95.55%
2.72%
80%
95.17%
2.56%
90%
94.27%
3.15%
95%
92.23%
4.02%
97%
87.97%
4.14%
Table 3. Recognition accuracy with IMU plus audio under different masking methods and factors. Mean per-activity accuracy over 20 activities and standard deviation across held-out participants.
Figure 5. Confusion matrices for four settings, rows are true labels and columns are predictions, 20 activities grouped as in Section 5.1 : (A) IMU only with trigger-selected segments; (B) audio only, no mask, trigger-selected segments; (C) IMU plus audio with a 10% point-wise mask; (D) IMU plus audio with a 50% contiguous mask. The label shut denotes the sound produced when closing the laptop.
Figure 6. t-SNE plots of KMeans clusters over fused features. (A) The 20 studied activities. (B) After adding the drawer activity. (C) After also adding the correction fluid activity. (D) After also adding the zipper activity. Each new activity appears as a separate cluster.
With each sensing modality exhibiting inherent strengths and limitations, multi-modal approaches for wearable Human Activity Recognition (HAR) are becoming increasingly relevant -- particularly for recognizing Activities of Daily Living (ADLs), where individual modalities often produce ambiguous signals for similar or complex activities. This work introduces HARMES, a multi-modal wearable dataset combining three wrist-recorded modalities: motion sensing via an Inertial Measurement Unit (IMU), atmospheric environmental sensors (humidity, temperature, and pressure), and audio. Collected from 20 participants performing household activities in their own homes, HARMES totals over 80 hours of recorded data, with approximately three hours of labeled activity data per participant across 15 ADL classes. To the best of our knowledge, HARMES is the first dataset to combine this particular sensor trio, and it is nearly six times larger than the previously largest wrist-inertial-acoustic HAR dataset. In an extensive benchmark, we evaluate cross-subject generalization and conduct an ablation study revealing that modality contributions are activity-dependent and can provide complementary value, particularly for activities that are ambiguous from motion data alone. HARMES is freely available at Zenodo, alongside example code for loading the dataset and training models on GitHub.
Robin Burchard, Pascal-André Brückner, Marius Bock +2
University of Siegen, Siegen, Germany · University of Bonn and & Lamarr Institute for Machine Learning and Artificial Intelligence, Bonn, Germany
Wearable human activity recognition (HAR) has made steady progress, yet much of this progress remains grounded in fixed-window, closed-set classification benchmarks. This formulation is poorly matched to everyday behavior, where activities are open-ended, unscripted, personalized, variable in duration, and often compositional. To address this mismatch, we introduce ActivityNarrated, an open-ended narrative paradigm for language-grounded wearable activity understanding. We formulate this setting as dense sensor signal captioning with a comprehensive benchmark protocol that measures temporal localization, caption quality, sensor-language alignment, conventional closed-set classification as a downstream diagnostic, and additional robustness measures. We further present ActNarrator, a 3-stage architecture that discretizes continuous IMU signals into reusable motion tokens and uses an external frozen small language model to generate open-vocabulary activity captions. Experiments show that our method provides high quality dense sensor captioning with superior adaptivity and robustness, enabling various downstream tasks by turning sensor-based human activity understanding into sensor-grounded text-level reasoning. This includes downstream classification where ActNarrator outperforms state-of-the-art HAR models by 3.8 - 31.6 % in Macro-F1. This paradigm also enables novel activity understanding capabilities such as complex question-answering over long time horizons.
Wearable human-activity recognition (HAR) models operate across sensors, subjects, and backbones, yet a smooth waveform may appear temporal while exploiting a persistent sensor offset primarily. We introduce SpectrumAudit, a label-sealed audit that fits a phase-randomized full-window stimulus on calibration windows from subjects held out from training and testing. After selection, it replays its exact DC projection and budget-constrained zero-mean residual on the same frozen victim without refitting. Across 27 victims from three datasets and three backbones, the selected waveforms cause 2.87-40.83-point three-phase robust accuracy losses. Under this replay budget, DC is more damaging than AC on 24/27 victims and recovers at least 90% of the full drop on 22/27; all 5 failures occur on WISDM. In a held-out UTD-MHAD check, the selected waveform causes 13.49-pp accuracy and 11.68-pp macro-F1 losses, versus -0.66 pp for matched random changes. The audit diagnoses offset versus zero-mean variation under a common peak-budget cap. The code will be released upon acceptance.
Qingyu Wu, Yuan Wei, Renju Liu +1
Defense Innovation Institute, Academy of Military Science, Beijing, China · School of Software Engineering, South China University of Technology, Guangzhou, China