GenGait: A Transformer-Based Model for Human Gait Anomaly Detection and Normative Twin Generation
Authors: Elisa Motta, Marta Lorenzini, Clara Mouawad, Alberto Ranavolo, Mariano Serrao, Arash Ajoudani
Organizations: HRI2 Laboratory, Istituto Italiano di Tecnologia (IIT), Genoa, Italy · Department of Occupational and Environmental Medicine, Epidemiology and Hygiene, INAIL, Rome, Italy · Department of Medical and Surgical Sciences and Biotechnologies, Sapienza University of Rome, Rome, Italy
Gait analysis provides an objective characterization of locomotor function and is widely used to support diagnosis and rehabilitation monitoring across neurological and orthopedic disorders. Deep learning has been increasingly applied to this domain, yet most approaches rely on supervised classifiers trained on disease-labeled data, limiting generalization to heterogeneous pathological presentations. The methodological objective of this work is to develop a label-free framework for joint-level anomaly detection and kinematic correction based on a Transformer masked autoencoder trained exclusively on normative gait sequences from 150 adults, acquired with a markerless multi-camera motion-capture system. At inference, a two-pass procedure is applied to potentially pathological input sequences: first, it estimates joint inconsistency scores by occluding individual joints and measuring deviations from the learned normative prior. Then, it withholds the flagged joints from the encoder input and reconstructs the full skeleton from the remaining spatiotemporal context, yielding corrected kinematic trajectories at the flagged positions. The validation objective is to assess whether the framework preserves unseen normative gait and reduces angular deviation in simulated abnormal gait patterns. In this proof-of-concept evaluation, data from 10 held-out normative participants, who performed seven simulated abnormal gait patterns, showed a significant reduction in angular deviation across all analyzed joints with large effect sizes, and preservation of normative kinematics. The proposed approach enables interpretable, subject-specific localization of joints that are inconsistent with learned normative gait patterns and generation of an individualized normative reconstruction without requiring disease labels. Video is available at https://youtu.be/Rcm3jqR5pN4.
Figures & tables
Figure 1: Pipeline overview. Five cameras at 30 Hz yield 3D joint positions. Pre-processing estimates missing joints via constrained interpolation and tokenizes sequences into J×T joint–frame tokens using a 7-frame sliding window with stride 1. Pass 1 gives the masking pattern producing mask m ; Pass 2 reconstructs masked joints using a MAE Transformer. Post-processing segments gait cycles via left heel-height peak detection.
Figure 2: Masked autoencoder Transformer for joint reconstruction (Pass 2). A 7-frame window is tokenized into joint–frame tokens and linearly projected, then indexed by a sinusoidal positional code P and three learned embeddings E (joint type, frame index, and motion/velocity). A mask pattern m (provided by Pass 1) specifies which token positions are hidden. The Token Masking operator replaces the selected tokens with a learned [MASK] placeholder, yielding a fixed-length masked sequence (visible tokens + [MASK] ) processed by an 8-layer Transformer encoder. The encoder outputs a full-length memory sequence; visible memory vectors are selected using memory∖m and injected into the decoder input via Token Assembly , while masked positions are filled with [MASK] ; P and E are added again before the 2-layer decoder. The Transformer decoder reconstructs the full token window, from which the reconstructed last-frame tokens are retained as the corrected pose estimate at inference.
Figure 3: Pass 1: mask identification. Training uses a curriculum of synthetic masks (random → structured) with temporally coherent spans to produce the mask list m . Inference uses tiled occlusions and a badness score Bj to select unreliable joints and produce m . In both cases m is then inputted to Pass 2 for masked reconstruction.
Figure 4: Skeletal reconstruction for the normative trial and the seven simulated anomalies. Each panel shows a different participant. Joints flagged as biomechanically inconsistent by Pass 1 are highlighted in red. Reconstructed skeletons (blue) are overlaid with input ones (gray). Participant IDs and flagged joints are labeled above each panel. Video animations for all conditions are available at https://youtu.be/GenGait .
Figure 5: Joint-angle trajectories across the normalized gait cycle for the normative trial and representative executions of the seven simulated abnormal-gait tasks. The panels show pelvis flexion/extension (F/E), right hip abduction/adduction (A/A), right hip F/E, and right knee F/E for selected participant–task pairs. Blue trajectories and shaded bands represent the normative reference ( μ±2σ ), original mean trajectories are shown in red, and reconstructed mean trajectories are shown in green. The horizontal axis represents the normalized gait cycle. The boundaries at 0% and 100% approximately delimit one right heel-strike-to-right-heel-strike cycle.
Participant
Trial
Pelvis flex/ext
Hip abd/add
Hip flex/ext
Knee flex/ext
Orig.
Recon.
Orig.
Recon.
Orig.
Recon.
Orig.
Recon.
mean ± SD
N
1.80±1.48
1.50±1.03
2.71±1.26
2.39±0.88
4.62±2.01
3.80±1.69
3.91±1.04
4.16±0.83
p007
CD
3.47
0.87
3.81
2.35
4.65
2.92
14.39
3.32
p003
HH
5.81
1.28
8.07
3.50
7.01
3.81
7.03
9.35
p004
HS
4.21
1.71
3.28
2.44
8.07
7.10
5.39
10.65
p010
GG
19.85
6.02
1.29
1.50
20.51
17.53
7.35
7.37
Table 1: RMSE (degrees) between observed gait and the normative reference for original simulated abnormal-gait trials and their reconstructions, across four joint angles. The normative row reports mean ± standard deviation across the 10 held-out normative participants. Each subsequent row corresponds to one representative participant selected for that instructed task.
Dataset
Pelvis F/E
Right hip A/A
Right hip F/E
Right knee F/E
Full analysis
−0.96 (95.7%)
−0.77 (82.9%)
−0.80 (84.3%)
−0.41 (61.4%)
Without p001
−0.95 (95.2%)
−0.82 (84.1%)
−0.81 (84.1%)
−0.39 (60.3%)
Without p002
−0.97 (96.8%)
−0.79 (84.1%)
−0.79 (82.5%)
−0.39 (58.7%)
Without p003
−0.96 (95.2%)
−0.79 (84.1%)
−0.79 (84.1%)
−0.43 (63.5%)
Without p004
−0.95 (95.2%)
−0.78 (84.1%)
−0.79 (84.1%)
−0.43 (61.9%)
Without p005
−0.95 (95.2%)
−0.72 (81.0%)
−0.83 (85.7%)
−0.42 (61.9%)
Table 2: Leave-one-participant-out sensitivity analysis. Each cell reports the rank-biserial correlation rrb , followed in parentheses by the percentage of the remaining participant–task pairs showing an RMSE reduction. Negative correlations indicate lower RMSE after reconstruction. The first row reports the original analysis of all 70 pairs; each subsequent row reports the analysis after excluding the seven pairs associated with the indicated participant. All leave-one-participant-out comparisons remained significant after Holm-Bonferroni correction ( pHolm<0.05 ).
Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these methods face challenges when foundational assumptions - such as consistent object scale, viewpoint, background, illumination, and centered placement - are violated. Those variations that occur render anomaly detection methods unusable in many real-world scenarios. To address these limitations, we introduce three key contributions: (1) a visual prompting pipeline that isolates objects using foreground-background masking; (2) a mechanism for unfreezing the teacher in student-teacher models to improve domain adaptability; and (3) a data augmentation strategy leveraging diffusion-generated synthetic images to enhance anomaly detection performance. We achieve a 3.5 percentage point improvement over the previous state-of-the-art on the challenging AeBAD dataset by using the Masked Multiscale Reconstruction (MMR) model as our backbone.
Mateo Diaz-Bone, Daniel Caraballo, Florian Scheidegger +11
Reconstruction errors in multivariate time-series anomaly detection may not reliably distinguish abnormal behavior from benign deviations. Language-derived semantics offer complementary context, but existing multimodal approaches may rely on time-associated paired textual information that is difficult to obtain consistently and is not provided by standard multivariate time-series anomaly detection benchmarks. This setting poses two challenges: (1) conditioning masked reconstruction on window-specific semantics without exposing exact numerical targets or anomaly-specific cues, and (2) using a window-independent concept of normality as a complementary semantic reference rather than an independent anomaly detector. We propose LLM-Enhanced Alignment and Reconstruction with Normality Guidance for Time Series (LEARN-TS), which uses a frozen language model to construct two role-separated semantic representations without requiring temporally paired external text. Window-specific observation semantics encode temporal and cross-variable context without exact numerical values to guide channel-shared patch-masked reconstruction. A fixed, dataset-agnostic normality prompt provides a window-independent semantic reference for aligning normal representations and estimating normality discrepancy. At inference, masking each temporal patch once yields timestamp-level reconstruction evidence, conditionally modulated by discrepancy from a separate unmasked view. Across four benchmarks, LEARN-TS achieves the highest mean performance in 13 of 16 dataset-metric comparisons. Controlled ablations examine observation conditioning, joint normality alignment and scoring, and reference content, showing dataset-dependent ranking benefits and modest average gains from semantic over random references.
Jahyeob Koo, Kio Yun, Byoungmo Koo +1
Department of Industrial and Management Engineering, Korea University Seoul, Republic of Korea
Industrial anomaly detection benefits from anomaly samples, yet newly deployed products typically provide only normal images, making anomaly samples difficult to collect. Zero-shot anomaly generation offers a promising solution which avoids collection of target-product anomalies. However, existing methods mainly rely on texture images or text descriptions as anomaly sources, which often produce unrealistic anomalies. Observing that similar anomalies can recur across different products, we propose anomaly transfer-based zero-shot generation, which reuses real anomalies from existing source products, making target-product anomalies no longer necessary to generate realistic anomalious samples for unseen target products. Since not every anomaly type suits the target product, an anomaly type filtering mechanism first selects plausible source types. To transfer selected anomaly, we propose DPA, a diffusion-based framework that decouples product-agnostic anomaly representations. Instead of directly extracting anomaly representations, DPA learns product-irrelevant anomaly embeddings through training with the mismatched data pair, enabling transferable anomaly concept learning across products. Furthermore, we design an adaptive mask-guided pipeline that leverages adaptive masks to control the positional and geometric plausibility of generated anomalies during generation. A training-free anomaly labeling module is further introduced to produce pixel-level annotations aligned with generated anomalies. Extensive experiments on MVTec-AD, VisA, and a dedicated anomaly-transfer benchmark demonstrate that the proposed setting and DPA generate more realistic anomalies and significantly improve downstream anomaly detection performance under both zero-shot and few-shot settings. Source code and models will be released.