Surface electromyography (sEMG) provides a wearable, noninvasive interface to neuromuscular activity for movement decoding and human-computer interaction. Population-scale decoding remains difficult because the relationship between sEMG and neuromuscular activity varies across users and sessions, while task-relevant dynamics span channels and multiple timescales. Learning waveform-to-output mappings from task labels leaves the distinction between recording variability and coordinated motor activity implicit. We introduce the Neuromotor Hierarchy Network (NHN), which learns a compact latent neuromotor state from task supervision to represent task-relevant neuromuscular coordination. NHN constructs this latent state through a hierarchy inspired by neuromotor organization. It adapts recording statistics while preserving relative intensity. Its spatiotemporal encoder uses parameter-efficient channel interactions and modulates features with multi-timescale history. The resulting features yield candidate activations of learned motor primitives, which are temporally integrated and continuously weighted to form the state. Theoretical analysis characterizes the efficiency, temporal behavior, and optimization of NHN's core mechanisms. We evaluate the architecture for both continuous hand-pose estimation on emg2pose and touch-typing recognition on emg2qwerty. On emg2pose, NHN reduces user-averaged angular error by 0.52% to 2.84% across all three generalization splits in both Regression and Tracking relative to Hadidi et al.'s best task-specific variants, using 48.42% to 48.51% fewer parameters. On emg2qwerty, NHN reduces beam-search character error rate by 19.40% zero-shot and 30.42% after fine-tuning relative to SplashNet-Upscale, using 65.86% fewer parameters. Physiology-guided inference of a latent neuromotor state supports parameter-efficient sEMG decoding.
Figures & tables
Figure 1: Physiological inspiration. Neuromotor coordination links hand motion and sEMG, motivating a compact latent neuromotor state.
Figure 2: NHN architecture. Task-trained NHN adapts and encodes sEMG, then integrates and allocates candidate drives to form a compact latent neuromotor state for pose and typing decoding.
User
Stage
User × Stage
Session
Efficiency
Model
AE ↓ ( ∘ )
LD ↓ (mm)
AE ↓ ( ∘ )
LD ↓ (mm)
AE ↓ ( ∘ )
LD ↓ (mm)
AE ↓ ( ∘ )
LD ↓ (mm)
Params. ( 106 ) ↓
GFLOPs per 5 s ↓
(a) Regression
Static a
16.86 ± 1.80
25.17 ± 2.84
19.18 ± 1.71
28.57 ± 2.46
18.87 ± 2.01
28.77 ± 2.78
17.87
26.81
0
0
Sensing Dynamics b
15.50 ± 1.40
21.80 ± 2.10
18.80 ± 1.60
26.60 ± 2.00
18.70 ± 1.60
27.20 ± 2.00
–
–
–
–
NeuroPose c
13.20 ± 1.10
17.50 ± 1.30
17.20 ± 1.70
24.00 ± 2.10
17.50 ± 1.50
24.90 ± 1.70
15.52
21.61
6.35
1.12
emg2pose d
12.57 ± 1.30
16.28 ± 1.81
15.17 ± 1.59
20.53 ± 2.13
15.57 ± 1.30
21.49 ± 1.71
14.01
18.94
2.95
2.21
Table 1: emg2pose Regression and Tracking. Split columns report user mean ± SD. Session columns weight sessions within each split, then splits equally. Bold marks column minima within each setting, excluding Static for efficiency. Dashes denote unreported values. Published split results: a Hadidi et al. (2026) , b,c,d Salter et al. (2024) .
ODV
TDV
TDT
Efficiency
Model
Greedy ↓ (%)
Beam ↓ (%)
Greedy ↓ (%)
Beam ↓ (%)
Greedy ↓ (%)
Beam ↓ (%)
Params. ↓ ( 106 )
GFLOPs ↓ per 30 s
(a) Zero-shot
TDS ConvNet a
72.44
72.07
55.57
52.10 ± 5.54
55.38
51.78 ± 4.61
5.29
61.61
SplashNet (Split-only) b
61.74
58.64
45.73
37.28 ± 6.91
45.69
37.37 ± 7.34
2.68
36.84
SplashNet-mini (Shared) b
61.07
58.20
45.33
36.46 ± 7.09
45.26
36.41 ± 7.30
1.38
36.84
SplashNet-Upscale b
60.16
56.95
44.79
35.49 ± 7.56
44.78
35.67 ± 6.79
2.58
71.38
Table 2: emg2qwerty CER (%). NHN reports mean ± SD across participants. Bold marks column minima within each setting. Dashes denote unreported results. ODV is not evaluated after fine-tuning. a Sivakumar et al. (2024) (ODV and efficiency from b ). b Hadidi et al. (2025) . c Mehlman et al. (2025) .
Figure 3: Learned organization of the compact latent neuromotor state. Pose Regression: (a) participant-mean Pearson correlations between relative primitive activations and joint angles. Dots mark ≥80% participant agreement with the mean sign. Dark/light strips indicate abduction-adduction/flexion-extension. Th/In/Mi/Ri/Li denote thumb/index/middle/ring/little fingers. (b) Initial and post-training adaptive-gain timescales ( τm , Equation 6 ) and mean integration lags, kernel-weighted mean ages of candidate drives (Equation 9 ), with five-run mean ± SD. S, M, and L denote short, intermediate, and long timescales. (c) Mean within-frame ranked allocation weights, with dark interquartile range (IQR) and light 5-95% frame bands. Hard-12/Uniform assign equal weights to 12/32 primitives. Typing: (d) within-participant cross-window key-profile cosine similarity. Circles denote participant-model observations, diamonds means, bars participant SD. (e) Initial and post-training adaptive-gain timescales and mean integration lags (five-run mean ± SD). (f) Ranked allocation weights with bands and reference distributions as in (c).
Finger-specific motor intent is a clinically meaningful control signal for post-stroke neurorehabilitation, where residual muscle activity may remain measurable despite weak or incomplete movement. We study five-finger multilabel intent decoding from impaired-arm high-density surface electromyography (sEMG) in PhysioMio, a bilateral longitudinal dataset collected from stroke patients. A common processing protocol aligns movement labels, applies 20--450 Hz Butterworth filtering and Symlet-4 wavelet denoising, segments overlapping 200 ms windows, and extracts twelve time- and frequency-domain descriptors per channel. Direct LSTM, CNN, and GNN baselines reveal complementary behavior: the LSTM attains the highest subset accuracy (0.545), whereas the GNN attains the highest macro F1 (0.706) and macro AUPRC (0.776). Architecture search then identifies CNN-Large as the strongest single-split CNN, with 0.593 subset accuracy and 0.714 macro F1, while CNN-Micro provides a compact architecture for embedded inference. To match a four-sensor hardware design, we retrain CNN-Micro using channels associated with ECRB, ECRL, FDS, and FDP and exclude the ground electrode from model input. Across five seeds, cross-channel knowledge distillation improves the four-channel student over direct training, reaching 0.5219±0.0114 subset accuracy, 0.7612±0.0038 finger accuracy, and 0.6095±0.0058 macro F1. The selected 123K-parameter model accepts nine windows of 48 features and has been exported to ONNX. These results establish a reproducible software path from post-stroke sEMG to compact five-finger intent prediction for subsequent hardware-in-the-loop evaluation.
Zakariyya Brewster, Divy Wadhwani, Emily Yan +5
Department of Engineering Science University of Toronto · Computer and Mathematical Sciences University of Toronto Scarborough · Department of Computer Science University of Toronto +3
Motor unit parameters such as the innervation zone centre or the conduction velocity of the electrical potential harbour the potential to improve the fidelity of neuromechanical models used for movement and force prediction. Determining these parameters in a non-invasive way is challenging, as they are subject-specific and may vary with muscle contraction. Existing work on the estimation of motor unit parameters mainly relies on white-box modelling and therefore requires substantial manual modelling effort. This work targets the simultaneous estimation of multiple subject-specific motor unit parameters from electromyography (EMG) recordings measured non-invasively at the skin surface. This results in an inverse problem with a nonlinear loss function. To address this problem, an informed autoencoder is developed. This autoencoder reconstructs the surface EMG recordings while learning the parameters in its latent space and adhering to physical laws that relate the parameters to the EMG signals. In experiments on synthetic data, innervation zone centres are estimated with a mean absolute error of 2.5989 mm, and conduction velocities of the electric potential are estimated with a mean absolute error of 0.1697 ms−1. These results demonstrate the plausibility of this novel approach, which enables the simultaneous estimation of several motor unit parameters while reducing manual modelling effort through the integration of data-driven machine learning.
Kaja Balzereit, Malte Mechtenberg, Axel Schneider
Hochschule Bielefeld, University of Applied Sciences and Arts, Institute for System Dynamics
Surface electromyography provides a practical way to infer human movement intention from wearable muscle recordings, but models trained under a single acquisition setting often lose reliability when the user, session, electrode layout, or gesture protocol changes. This paper proposes AEMG, a self-supervised learning approach designed to extract reusable neuromuscular representations from diverse EMG sources. Eight public gesture datasets are first transformed into a shared signal format to reduce discrepancies in channel configuration, sensor topology, and recording protocol. Instead of relying on fixed-length sliding windows, AEMG identifies contraction events from energy variations and represents them as compact neuromuscular tokens, while ordered token groups describe the coordinated activity of multiple muscles during motion. A spatially and temporally conditioned Transformer is then used to encode these token sequences, preserving information about electrode position, activation timing, and sequential structure. For pre-training, the model constructs a discrete library of contraction prototypes through vector-quantized reconstruction and further learns contextual dependencies by recovering masked neuromuscular tokens from surrounding observations. Experiments under leave-one-subject-out and low-label adaptation settings show that the learned representation improves robustness to unseen users and reduces the amount of calibration data required for gesture recognition. These findings suggest that event-level token modeling offers a scalable route toward adaptable and data-efficient EMG-based motor-intent understanding.