cs.LGSep 30, 2026

EFormer: Temporally Aligned Local Correction for Continuous sEMG-Based Hand Pose Tracking

Authors: JiaCheng Ge, SiYu Zhang

Abstract

Surface electromyography (sEMG) provides a wearable, camera-free signal for continuous hand-motion inference. Mapping muscle activity to joint kinematics remains challenging because the recorded waveforms are indirect measurements, their relationship with motion changes over time, and individual anatomy and sensor placement alter the signal distribution. This paper presents EFormer, a residual feature-correction network built on a frozen tracking backbone. EFormer combines a high-rate event branch, temporally aligned local cross-attention, two causal rotary position embedding (RoPE) temporal layers, and a bounded, dynamically gated residual. EFormer receives 16-channel sEMG sampled at 2 kHz and fuses a 64-channel tracking representation at 25 Hz with a 128-channel event representation at 200 Hz. Cross-attention uses a nominal delay of 100 ms, a 300 ms history parameter, and a 50 ms tolerance; its causal mask restricts each query to events occurring 50-400 ms earlier. The correction scale is 0.15. The evaluated continuation-training configuration contains 585,376 trainable parameters and 5,974,508 frozen parameters. On the test set, EFormer achieves an MAE of 0.1546634 rad, an RMSE of 0.24063 rad, and an R^2 of 0.74801, compared with 0.1745326 rad, 0.2715448 rad, and 0.6791103 for the official tracking baseline. EFormer reduces MAE by 11.38% relative to the baseline. The results show that temporally aligned event-feature correction can reduce continuous hand-pose tracking error.

Explore similar work

Oct 4, 2026cs.LG

EMG-GPT: Predictive Pretraining on Residual-Quantized EMG Tokens for Hand Pose Estimation

Surface electromyography (sEMG) is a low-power, cost-effective biosignal for hand-pose estimation and gesture classification. In this work, we examine whether self-supervised pretraining on sEMG can yield transferable representations for continuous hand-pose estimation. We introduce EMG-GPT, a causal transformer-based model that operates on discrete sEMG representations from a frozen residual vector quantization (RVQ) tokenizer and learns temporal dynamics through depth-autoregressive future-code prediction. The model combines within-frame integration with causal temporal modeling while preserving the geometry of the pretrained codebook. EMG-GPT shows competitive results in both Regression and Tracking tasks, supporting EMG-only pretraining as a viable approach for learning transferable sEMG representations.
May 7, 2026cs.CV

EgoEMG: A Multimodal Egocentric Dataset with Bilateral EMG and Vision for Hand Pose Estimation

Surface electromyography (sEMG) records muscle activity during hand movement and can be decoded to recover detailed hand articulation. EMG and egocentric vision are complementary for hand sensing: EMG captures fine-grained finger articulation even under occlusion and poor lighting, while vision provides global hand configuration. However, no existing dataset synchronizes both modalities. We present EgoEMG, a multimodal egocentric dataset for bimanual hand pose estimation. EgoEMG includes bilateral wristband EMG with 16 total channels (8 per wrist) sampled at 2 kHz, 120 Hz IMU, egocentric wide-angle RGB video, external RGB-D video, and mocap-derived hand motion with wrist articulation angles. The dataset covers 41 participants performing 60 gesture classes, including 30 single-hand gestures and 30 bimanual gestures, totaling more than 10 hours of recording. We also introduce a benchmark with three tasks -- EMG-to-pose, vision-to-pose, and EMG+vision fusion -- under a shared joint-angle prediction target and common generalization split axes (cross-gesture, cross-user, and combined). As baselines, we evaluate EMGFormer for EMG-to-pose and generic ResNet/ViT backbones for vision-to-pose. We further study a residual fusion architecture that improves over matched lightweight vision-only baselines. Together, EgoEMG and its benchmark establish a foundation for future research on multimodal hand pose estimation with EMG and vision.
Apr 24, 2026cs.LG

Decoding High-Dimensional Finger Motion from EMG Using Riemannian Features and RNNs

Continuous estimation of high-dimensional finger kinematics from forearm surface electromyography (EMG) could enable natural control for hand prostheses, AR/XR interfaces, and teleoperation. However, the complexity of human hand gestures and the entanglement of forearm muscles make accurate recognition intrinsically challenging. Existing approaches typically reduce task complexity by relying on classification-based machine learning, limiting the controllable degrees of freedom and compromising on natural interaction. We present an end-to-end framework for continuous EMG-to-kinematics regression using only consumer-grade hardware. The framework combines an 8-channel EMG armband, a single webcam, and an automatic synchronization procedure, enabling the collection of the EMG Finger-Kinematics dataset (EMG-FK), a 10-h dataset of synchronized EMG and 15 finger joint angles from 20 participants performing rich, unconstrained right-hand motions. We also introduce the Temporal Riemannian Regressor (TRR), a lightweight GRU-based model that uses sequences of multi-band Riemannian covariance features to decode finger motion. Across EMG-FK and the public emg2pose benchmark, TRR outperforms state-of-the-art methods in both intra- and cross-subject evaluation. On EMG-FK, it reaches an average absolute error of 9.79°±1.489.79 °\pm 1.48 in intra-subject and 16.71°±3.9716.71 °\pm 3.97 in cross-subject. Finally, we demonstrate real-time deployment on a Raspberry Pi 5 and intuitive control of a robotic hand; TRR runs at nearly 10 predictions/s and is roughly an order of magnitude faster than state-of-the-art approaches. Together, these contributions lower the barrier to reproducible, real-time EMG-based decoding of high-dimensional finger motion, and pave the way toward more natural and intuitive control of embedded EMG-based systems.