MorphCL: Morphological Contrastive Learning for Inertial-based Human Activity Recognition
Organizations: University of Bonn · University of Cambridge · Lamarr Institute for Machine Learning and Artificial Intelligence · University of Siegen
Abstract
Despite the ubiquity of sensors in wearable and mobile devices and the abundance of human movement data they generate, translating unlabeled recordings into foundational motion models remains an open challenge. Self-supervised learning (SSL) has alleviated the need for costly annotations, yet existing approaches leave the global structure of large-scale motion data largely untapped, relying on randomly sampled batches and local comparisons that become particularly problematic for in-the-wild inertial data dominated by stationary, low-variance behaviors. Here we introduce Morphological Contrastive Learning (MorphCL), a self-supervised pretraining framework that uses structure-aware grouping to inject explicit modeling of global structure into inertial-based SSL approaches. Building on two well-established pillars of motion analysis, the discovery of motion primitives, or motifs, and domain-specific feature descriptors, we show that MorphCL substantially improves linear probing and finetuning results of learned encoders by up to 15 percentage points in F1-score. In a comparison with existing foundation models, we demonstrate that MorphCL-pretrained encoders match or surpass them models in linear probing performance while trained on less data. Qualitative analysis of the resulting embedding spaces further reveals morphologically meaningful cluster structure, with improved separation of kinematically similar activity classes.
Figures & tables
| Dataset | Scenario | Activity Types | #Sbjs | #Cls | Avg. Recording (min) |
| Bock et al. (2024) | outdoor sports workouts | periodic, complex | 22 | 18(+1) | 52.52 ( 17.89) |
| Scholl et al. (2015) | wetlab DNA extraction | (non-)periodic | 22 | 8(+1) | 47.93 ( 7.48) |
| Sztyler & Stuckenschmidt (2016) | locomotion | periodic | 13 | 8 | 75.84 ( 4.55) |
| Hoelzemann et al. (2023) | basketball practice and game | (non-)periodic | 24 | 5(+1) | 60.59 ( 10.84) |
| Reiss & Stricker (2012) | locomotion and ADL | periodic, complex | 9 | 18(+1) | 50.46 ( 17.79) |
| DeepConvLSTM WISDM | ViT-S WISDM | ViT-S CAPTURE-24 | |||||||
| Method | Base | + MorphCL | Base | + MorphCL | Base | + MorphCL | |||
| MTL | |||||||||
| SimCLR | |||||||||
| Masked AE | – | – | – | ||||||
| Denoising AE | – | – | – | ||||||
| REBAR | |||||||||
| DeepConvLSTM WISDM | ViT-S WISDM | ViT-S CAPTURE-24 | |||||||
| Method | Base | + MorphCL | Base | + MorphCL | Base | + MorphCL | |||
| MTL | |||||||||
| SimCLR | |||||||||
| Masked AE | – | – | – | ||||||
| Denoising AE | – | – | – | ||||||
| REBAR | |||||||||
| MTL | SimCLR | Masked AE | |||||||
| Subset | Base | + MorphCL | Base | + MorphCL | Base | + MorphCL | |||
| Full | |||||||||
| Motif | |||||||||
| Morph | |||||||||
| Hours of Data | Model Size | WEAR | Wetlab | PAMAP2 | Hangtime | RWHAR | Avg. | |||||||
| LProbe | Fine | LProbe | Fine | LProbe | Fine | LProbe | Fine | LProbe | Fine | LProbe | Fine | |||
| Yuan et al. (2024) - MTL | ||||||||||||||
| ResNet18 | 3.7K | 10.4M | ||||||||||||
| ResNet18 | 16.8M | 10.4M | ||||||||||||
| Ours - MTL + MorphCL | ||||||||||||||
| ResNet18 | 3.7K | 10.4M | ||||||||||||
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Pretraining Corpus | Unfiltered | Motif-filtered | Morph-filtered | |||
| Train | Test | Train | Test | Train | Test | |
| WISDM | 35.5 | 9.7 | 15.2 | 4.4 | 3.2 | 0.8 |
| CAPTURE-24 | 3004.6 | 787.4 | 783.5 | 197.7 | 47.3 | 39.0 |
| WEAR | Wetlab | PAMAP2 | Hangtime | RWHAR | Avg. | |||||||
| LProbe | Fine | LProbe | Fine | LProbe | Fine | LProbe | Fine | LProbe | Fine | LProbe | Fine | |
| Encoder: DeepConvLSTM | ||||||||||||
| Fully-supervised | ||||||||||||
| Pretraining: WISDM | ||||||||||||
| MTL | ||||||||||||
| SimCLR | ||||||||||||
| WEAR | Wetlab | PAMAP2 | Hangtime | RWHAR | Avg. | |
| Encoder: DeepConvLSTM; Pretraining: WISDM | ||||||
| MTL | ||||||
| MTL + MorphCL | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| SimCLR | ||||||
| SimCLR + MorphCL | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| REBAR | ||||||
| WEAR | Wetlab | PAMAP2 | Hangtime | RWHAR | Avg. | |
| Encoder: DeepConvLSTM; Pretraining: WISDM | ||||||
| MTL | ||||||
| MTL + MorphCL | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| SimCLR | ||||||
| SimCLR + MorphCL | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| REBAR | ||||||
| Discovered motifs | Extracted windows | Relative to full dataset | Downstream F1 | |
| Discovered motifs | Extracted windows | Relative to full dataset | Downstream F1 | |
| Noise | DBCV | AMI (vs. activity) | Downstream F1 | ||
| – | – | – | – |
| Support | Noise | DBCV | AMI (vs. activity) | Downstream F1 | |
| Method | Full | Motif | Morph | Random ( =Motif) | Random ( =Morph) |
| SimCLR | |||||
| SimCLR + MorphCL | |||||
| MTL | |||||
| MTL + MorphCL | |||||
| Denoising AE | |||||
| Denoising AE + MorphCL |
| Method | Full | Motif | Morph | Random ( =Motif) | Random ( =Morph) |
| SimCLR | |||||
| SimCLR + MorphCL | |||||
| MTL | |||||
| MTL + MorphCL | |||||
| Denoising AE | |||||
| Denoising AE + MorphCL |
| Method | Balanced + Oversampling | Balanced + Undersampling | Unbalanced |
| SimCLR + MorphCL | |||
| MTL + MorphCL |
| Method | Balanced + Oversampling | Balanced + Undersampling | Unbalanced |
| SimCLR + MorphCL | |||
| MTL + MorphCL |
| Framework | WEAR | Wetlab | PAMAP2 | Hangtime | RWHAR | WISDM | Avg |
| Encoder: DeepConvLSTM | |||||||
| Pretraining: WISDM | |||||||
| MTL | -0.1776 | -0.2189 | 0.0533 | -0.0471 | -0.0639 | -0.2035 | -0.1096 |
| MTL + MorphCL | -0.2019 | -0.2025 | 0.1088 | -0.0493 | -0.0872 | -0.1682 | -0.1001 |
| SimCLR | -0.1990 | -0.2002 | 0.0449 | -0.0742 | -0.0928 | -0.1814 | -0.1171 |
| SimCLR + MorphCL | -0.1906 | -0.1964 | 0.0741 | -0.0780 | -0.1096 | -0.1961 | -0.1161 |
| WEAR | Wetlab | PAMAP2 | Hangtime | RWHAR | |
| Encoder: DeepConvLSTM; Pretraining: WISDM | |||||
| Khosla et al. (2020) | |||||
| Pillai et al. (2025) | |||||
| Ours | |||||
| Encoder: ViT-S; Pretraining: WISDM | |||||
| Khosla et al. (2020) | |||||