Surface electromyography (sEMG) is a low-power, cost-effective biosignal for hand-pose estimation and gesture classification. In this work, we examine whether self-supervised pretraining on sEMG can yield transferable representations for continuous hand-pose estimation. We introduce EMG-GPT, a causal transformer-based model that operates on discrete sEMG representations from a frozen residual vector quantization (RVQ) tokenizer and learns temporal dynamics through depth-autoregressive future-code prediction. The model combines within-frame integration with causal temporal modeling while preserving the geometry of the pretrained codebook. EMG-GPT shows competitive results in both Regression and Tracking tasks, supporting EMG-only pretraining as a viable approach for learning transferable sEMG representations.
Figures & tables
Figure 1: Overview of EMG-GPT. sEMG patches are mapped by frozen NeuroRVQ to structured time–electrode–branch–depth codes. Frozen vector sums are integrated within each frame and modeled causally across frames. The pre-head state either predicts a future frame coarse-to-fine with code-index cross-entropy, or bypasses the categorical head and transfers to parallel Regression and Tracking decoders under frozen or full-backbone adaptation. Regression interpolates features to the 50 -Hz pose grid without an initial pose, whereas Tracking repeats features and receives the boundary pose. Only token-frame modeling is causal.
Figure 2: Tracking poses for EMG-GPT and vEMG2Pose. Top: motion capture; middle: EMG-GPT Tracking; bottom: vEMG2Pose ( Salter et al., 2024 ) .
Test set
Model
Angular MAE ( ∘ ) ↓
Landmark (mm) ↓
User
SensingDynamics
15.5 ± 1.4
21.8 ± 2.1
NeuroPose
13.2 ± 1.1
17.5 ± 1.3
vEMG2Pose
12.2 ± 1.3
15.8 ± 1.9
EMG-GPT
12.8 ± 1.1
16.3 ± 1.3
Stage
SensingDynamics
18.8 ± 1.6
26.6 ± 2.0
NeuroPose
17.2 ± 1.7
24.0 ± 2.1
Table 1: Regression test results. Published baselines are from Salter et al. (2024) ; EMG-GPT is a fully adapted model. Angular and landmark values are mean ± sample SD across users within each generalization condition.
Table 2: Configuration of the selected 25-Hz, K=4 EMG-GPT model.
Test set
Model
Angular MAE ( ∘ ) ↓
Landmark (mm) ↓
User
vEMG2Pose
7.7 ± 1.0
10.3 ± 1.5
EMG-GPT
8.9 ± 0.9
11.6 ± 1.2
Stage
vEMG2Pose
11.2 ± 1.4
15.2 ± 1.9
EMG-GPT
12.4 ± 1.4
16.9 ± 1.8
User, Stage
vEMG2Pose
11.0 ± 1.0
15.4 ± 1.4
EMG-GPT
12.2 ± 1.3
16.9 ± 1.4
Appendix
Table 3: Tracking test results, where the boundary pose is given. vEMG2Pose values are from ( Salter et al., 2024 ) .
Test set
Model
Angular MAE ( ∘ ) ↓
Landmark (mm) ↓
Regression
User
vEMG2Pose
12.80 ± 1.31
16.29 ± 1.85
EMG-GPT
12.80 ± 1.08
16.30 ± 1.29
Stage
vEMG2Pose
15.40 ± 1.59
20.43 ± 2.09
EMG-GPT
16.30 ± 1.57
22.02 ± 1.92
User, Stage
vEMG2Pose
15.73 ± 1.38
21.25 ± 1.79
Appendix
Table 4: Locally matched test evaluation of the released vEMG2Pose checkpoints and EMG-GPT. Values are mean ± sample SD across users within each generalization condition. Bold indicates the lower mean; rounded ties are bolded together. Each method retains its native input frontend.
Figure 3: Median held-out user and median held-out stage ( FingerFreeform ). Top: motion capture; middle: EMG-GPT Tracking; bottom: vEMG2Pose. Six poses unroll from left to right over a five-second validation clip.
Figure 4: Median held-out user and lower-error held-out stage ( ShakaVulcanPeace ). Rows and temporal layout follow Figure 3 .
Figure 5: Median held-out user and higher-error held-out stage ( CountingUpDownFingerWigglingSpreading ). Rows and temporal layout follow Figure 3 .
Figure 6: Lower-error held-out user and median held-out stage ( HandOverHandCountingUpDownFingerWigglingSpreading ). Rows and temporal layout follow Figure 3 .
Figure 7: Higher-error held-out user and median held-out stage ( FingerFreeform ). Rows and temporal layout follow Figure 3 .
Figure 8: Regression example for a median held-out user and median held-out stage. Top: motion-capture target; middle: EMG-GPT Regression; bottom: vEMG2Pose Regression. Six poses unroll from left to right over the same target-aligned five-second validation clip.
Figure 9: Regression example for a median held-out user and lower-error held-out stage. Rows and target-aligned temporal layout follow Figure 8 .
Figure 10: Regression example for a median held-out user and higher-error held-out stage. Rows and target-aligned temporal layout follow Figure 8 .
Public surface electromyography (EMG) datasets vary widely in electrode layout, channel count, frequency support, and size. Simply mixing them for pretraining can misalign channel semantics, introduce spectral targets that some devices cannot observe, and let large or high-channel-count datasets dominate learning. We introduce EMGBlend, a self-supervised framework designed around these differences. It combines shared channel patches with geometry-aware attention, restricts spectral targets to each recording's supported frequency band, and balances exposure across data sources. We pretrain a 109M-parameter model on 11 public EMG sources and evaluate it on gesture recognition, continuous-force regression, and contact classification. EMGBlend consistently outperforms matched random initialization and waveform reconstruction controls. Fixed-budget source controls show that multi-source pretraining improves gesture recognition and remains competitive for force decoding. Ablations confirm that geometry, band-aware targets, and source balancing each contribute to transfer, although cross-person NinaPro force estimation remains difficult. Overall, EMGBlend shows how heterogeneous EMG datasets can be combined through explicit mechanism design rather than simple concatenation. Code is available at https://github.com/tamanano/EMGBlend
Yuwei Jia, Cheng Zhong, Jinyang Yu +1
Beijing University of Posts and Telecommunications, China · Dexwise, China · Shenzhen University, China
Surface electromyography provides a practical way to infer human movement intention from wearable muscle recordings, but models trained under a single acquisition setting often lose reliability when the user, session, electrode layout, or gesture protocol changes. This paper proposes AEMG, a self-supervised learning approach designed to extract reusable neuromuscular representations from diverse EMG sources. Eight public gesture datasets are first transformed into a shared signal format to reduce discrepancies in channel configuration, sensor topology, and recording protocol. Instead of relying on fixed-length sliding windows, AEMG identifies contraction events from energy variations and represents them as compact neuromuscular tokens, while ordered token groups describe the coordinated activity of multiple muscles during motion. A spatially and temporally conditioned Transformer is then used to encode these token sequences, preserving information about electrode position, activation timing, and sequential structure. For pre-training, the model constructs a discrete library of contraction prototypes through vector-quantized reconstruction and further learns contextual dependencies by recovering masked neuromuscular tokens from surrounding observations. Experiments under leave-one-subject-out and low-label adaptation settings show that the learned representation improves robustness to unseen users and reduces the amount of calibration data required for gesture recognition. These findings suggest that event-level token modeling offers a scalable route toward adaptable and data-efficient EMG-based motor-intent understanding.
Surface electromyography (sEMG) provides a wearable, camera-free signal for continuous hand-motion inference. Mapping muscle activity to joint kinematics remains challenging because the recorded waveforms are indirect measurements, their relationship with motion changes over time, and individual anatomy and sensor placement alter the signal distribution. This paper presents EFormer, a residual feature-correction network built on a frozen tracking backbone. EFormer combines a high-rate event branch, temporally aligned local cross-attention, two causal rotary position embedding (RoPE) temporal layers, and a bounded, dynamically gated residual. EFormer receives 16-channel sEMG sampled at 2 kHz and fuses a 64-channel tracking representation at 25 Hz with a 128-channel event representation at 200 Hz. Cross-attention uses a nominal delay of 100 ms, a 300 ms history parameter, and a 50 ms tolerance; its causal mask restricts each query to events occurring 50-400 ms earlier. The correction scale is 0.15. The evaluated continuation-training configuration contains 585,376 trainable parameters and 5,974,508 frozen parameters. On the test set, EFormer achieves an MAE of 0.1546634 rad, an RMSE of 0.24063 rad, and an R^2 of 0.74801, compared with 0.1745326 rad, 0.2715448 rad, and 0.6791103 for the official tracking baseline. EFormer reduces MAE by 11.38% relative to the baseline. The results show that temporally aligned event-feature correction can reduce continuous hand-pose tracking error.