cs.AIOct 5, 2026

Signature-Based Feature Learning for Human Activity Recognition: A Reproducible Machine Learning Study of Representation, Depth, and Model Choice

Authors: Kamal Jarrar, Jacky Cresson, Christian Paroissin

Organizations: LMAP

Abstract

Human activity recognition (HAR) relies on transforming sensor signals into informative representations for classification. Although deep learning and handcrafted features are widely used, the role of representation itself is often not systematically isolated. Signature transforms provide a mathematically grounded way to encode temporal order and cross-channel interactions, but their value for HAR under a fully reproducible and leakage-aware framework remains unclear. To evaluate whether signature-based feature learning improves HAR performance compared with raw-signal baselines, and to assess the effects of embedding strategy, truncation depth, model choice, and sensor configuration. Experiments were conducted on the UCI HAR dataset using a fully reproducible pipeline with the original train--test split preserved and subject-disjoint validation to prevent leakage. Three representations were compared: raw flattened signals, time-augmented paths, and lead--lag transformed paths. Signature features were computed at multiple truncation depths and evaluated using multilayer perceptron (MLP) and Random Forest (RF) classifiers under identical preprocessing and validation procedures. A prior K-means-based feature reduction study was also reproduced for comparison. Signature-based representations improved performance when paired with RF models, the best configuration was time-augmented six-channel signatures at depth 6 using entropy-based RF achieving 0.858 accuracy and 0.859 macro F1, outperforming the strongest raw baseline (0.816 accuracy). Lead--lag representations were competitive at moderate depths but did not surpass the best time-augmented models. MLP models did not exceed raw baselines. Signature-based feature learning can improve HAR, but its benefit depends on alignment between representation design and classifier choice.

Explore similar work

Jul 29, 2026cs.LG

RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies. Retrieval-Augmented Generation for Human Activity Recognition (RAG-HAR) addresses this by framing HAR as a training-free, retrieval-augmented task, in which statistical descriptions of sensor windows are used to retrieve similar labeled examples that guide LLM-based classification. We introduce RAG-HAR+, a retrieval-first and cost-optimized extension that strengthens retrieval while reducing dependence on LLM-based inference. RAG-HAR+ uses an offline Retrieval Designer Agent to design dataset-specific feature groups from a diverse pool of motion descriptors, enabling sensor windows to be compared using features better aligned with dataset-specific activity patterns. During inference, RAG-HAR+ uses majority voting over retrieved neighbors for samples with strong retrieval evidence and defers only uncertain cases to an LLM-based Ambiguity Resolver Agent. Across six HAR benchmarks, RAG-HAR+ maintains competitive or improved performance while reducing LLM usage, token consumption, and inference time. We further extend the RAG-HAR mobile prototype to demonstrate the practical feasibility of retrieval-first, LLM-assisted HAR in mobile sensing scenarios.
Sep 28, 2026cs.AI

More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with Jev

General-purpose models promise sensor-based decisions without training a task-specific classifier, which could reduce the dependence of Human Activity Recognition (HAR) on labeled data. Yet it remains unclear whether such models can directly interpret deterministic descriptions of physical sensor signals well enough to replace or complement trained HAR models. We study this question using Jev, a fixed general-purpose probabilistic decision model, on 1,800 class-balanced accelerometer windows from WISDM, UCI341, and PAMAP2. Jev receives no labeled examples, retrieval context, or HAR-specific parameter updates. We evaluate three deterministic sensor representations and compare 5,400 Jev decisions with a generative baseline and three supervised HAR models. Jev remains far below supervised recognition, with its strongest representation reaching macro-F1 of 0.038, 0.118, and 0.089 across the three datasets, compared with 0.686 to 0.907 for the supervised models. More numerical features do not improve Jev. Instead, they reduce recognition on all three datasets, while augmenting the same numerical evidence with a deterministic semantic rendering partially recovers performance, although the experiment does not isolate semantics from the accompanying serialization and redundancy changes. Jev is fast and inexpensive to query, but its probabilities are not reliably calibrated for recognition. A post-hoc fusion analysis finds a small improvement on WISDM that does not replicate on UCI341 or PAMAP2. These results show that training-free sensor decisions depend not only on the information available in the signal, but also on whether the model can use the representation through which that information is exposed. The sensor-to-model interface should therefore be treated as part of the model evaluation rather than as a neutral preprocessing step.
Jul 17, 2026eess.SP

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition

Sensor-based human activity recognition (HAR) has achieved significant progressed in fully supervised learning settings. However, these supervised learning models rely on large amount of labeled data, which require labor-intensive collection and meticulous annotation. To address these challenges, this paper proposes a Joint Embedding Predictive Architecture framework tailored for sensor-based HAR, designed to learn robust and generalizable representations from unlabeled datasets. The proposed framework features an encoder designed to explicitly model both the fine-grained local temporal representations within individual window and the long-term temporal sequence of adjacent windows. Furthermore, we introduce an improved Variance-Invariance-Covariance Regularization (VICReg) objective function that incorporates computationally lightweight norm term to stabilize the JEPA pre-training phase. This term balances variance, invariance and covariance constraints to prevent representation collapse. The proposed HAR-JEPA framework is evaluated using two benchmark continuously performed activity datasets. The results show that high-quality representations are successfully learned by the proposed framework. Furthermore, the representations learned by HAR-JEPA demonstrates superior generalization on minority, high variance transitional activities such as sit-to-stand and sit-to-lie where supervised learning tend to overfit due to limited support.