Deploying intelligent robotic systems that interact with humans through gestures requires neural networks capable of recognizing diverse temporal patterns. We present a systematic benchmark of ten abstract sequential tasks--five permutation-invariant (set) and five order-dependent (sequence) problems--evaluated across eighteen neural network architectures spanning recurrent, convolutional, attention-based, and set-function families. Beyond the core architecture-task grid, we explore numerous preprocessing and target-variable transformations, yielding more than 250 distinct experimental configurations. All variants are trained and tested under strictly identical conditions (fixed random seeds, shared hyperparameters, shared data splits) to ensure fair and reproducible comparison. Ranking across all ten tasks reveals four consistently top-performing architectures--BiGRU, TCN, Conv1D, and GRUReLU--all compact enough for real-time deployment (under 2,000 parameters in the benchmark setting). Based on this ranking, we apply three architecturally diverse top models (BiGRU, TCN, and GRUReLU) to a practical robotics problem: estimating the execution speed of dynamic arm gestures from skeletal keypoint sequences. Three speed interpretations (peak count, period time, and mean spike spacing) are evaluated on a custom dataset of eight traffic-related gesture classes comprising 256,710 frames recorded via OpenPose. The best configuration achieves a mean absolute error of 0.198 on the peak-count interpretation, corresponding to roughly 5% relative error, while the period-time interpretation reaches approximately 4% relative error, and the mean spike spacing interpretation approximately 8% relative error. These results demonstrate that neural networks can reliably estimate gesture speed from skeletal data, opening a path toward speed-aware gesture-controlled robotic systems.
Figures & tables
Architecture
Params
1-Decision
2-Counting
3-Variance
4-Mean
5-Max
GRU
14
1.000
3.428 11
0.004
0.001 2
0.006 4
GRUReLU
963
0.994 2
1.226 9
0.004
0.002 3
0.011 9
BiGRU
1,921
1.000
2.020 10
0.004
0.001 2
0.007 5
LSTM
16
0.960 5
12.726 14
0.025 4
0.012 7
0.020 11
BiLSTM
2,529
0.988 4
14.126 15
0.012 3
0.015 9
0.033 13
SkipGRU
14
1.000
19.444 16
0.062 9
0.014 8
0.044 14
Table 1. Performance on set tasks ( 1000:1000 configuration). Task 1: accuracy ( ↑ ); Tasks 2–5: MAE ( ↓ ). Values rounded to 3 decimal places. Bold = rank 1. Superscript = per-task dense rank (ties share the same rank). – = convergence failure (ranked last).
Architecture
Params
6-Speed
7-Peak
8-Period
9-SpikeSp
10-Threshold
GRU
14
0.004
225.496 18
2.500 13
6.752 2
10.318 13
GRUReLU
963
0.006 3
123.024 11
0.074 4
6.752 2
6.294 2
BiGRU
1,921
0.008 4
8.023 3
0.031 2
6.752 2
6.243
LSTM
16
0.083 9
225.267 17
2.500 13
6.753 3
133.807 16
BiLSTM
2,529
0.069 7
113.481 10
0.092 5
6.752 2
6.302 3
SkipGRU
14
0.063 6
180.412 15
2.501 14
6.759 6
241.764 17
Table 2. Performance on sequence tasks ( 1000:1000 configuration, MAE ↓ ). Values rounded to 3 decimal places. Bold = rank 1. Superscript = per-task dense rank (ties share the same rank). – = convergence failure (ranked last).
#
Model
Params
Σrank
ΣSET
ΣSEQ
1
BiGRU
1,921
31
19
12
1
TCN
1,667
31
19
12
2
Conv1D
292
43
20
23
3
GRUReLU
963
46
24
22
4
Transformer
24,961
54
19
35
5
Quantile
385
59
21
38
Table 3. All architectures ranked by global rank sum ( Σrank , lower is better). ΣSET and ΣSEQ show rank sums for set and sequence tasks, respectively.
Figure 1. Per-task dense rank of all 18 architectures across the 10 benchmark tasks ( 1000:1000 configuration). Each point shows one architecture’s rank on a single task; legend entries are ordered by global rank sum ( Σ ). Blue background: permutation-invariant set tasks (Tasks 1–5); orange: order-dependent sequence tasks (Tasks 6–10). Bump chart with one colored series per architecture, showing each model's per-task rank across the ten benchmark tasks.
Figure 2. Cross-task performance profiles of the five architectures with the lowest global rank sum ( 1000:1000 configuration). Spoke values are min–max normalized across all 18 architectures (1.0 = best, 0.0 = worst). Blue sector: set tasks; orange sector: sequence tasks. Radar chart showing ten task spokes; each of the five top-ranked architectures is drawn as a colored polygon indicating its normalized score on every task.
Figure 3. Reference poses for each of the eight gesture types, visualized as BODY-25 skeletal keypoints after 1×1 normalization. Blue: right arm; red: left arm; black: torso; green: head. Eight skeleton plots showing the normalized reference position for each gesture type used in the speed estimation application.
Figure 4. Weighted distance-from-reference-pose curves for example gesture sequences. Red dots mark local minima—moments at which the skeleton is closest to the reference pose. Inter-minima spacing reflects execution speed. Line plots showing the weighted distance from the reference pose over time for different gestures, with local minima marked by red dots.
Interpr.
Model
Params
MAE
MSE
Bias
Std.
7-Peak
BiGRU
27,649
0.198
0.097
+ 0.009
0.311
GRUReLU
13,825
0.214
0.113
− 0.009
0.336
TCN
6,785
0.272
0.141
+ 0.065
0.370
8-Period
GRUReLU
13,057
1.320
5.942
− 0.220
2.428
BiGRU
26,113
1.499
7.157
+ 0.478
2.632
TCN
6,401
2.882
16.756
+ 0.775
4.019
Table 4. Application results on real gesture speed estimation. Bold = best per interpretation group: MAE, MSE, Std ( ↓ ); Params ( ↓ fewest); Bias ( ∣⋅∣ closest to zero). Bias is the signed mean error; Std is the error standard deviation.
Figure 5. MAE and MSE for the three selected architectures across the three speed interpretations. Grouped bar chart comparing the MAE of BiGRU, GRUReLU, and TCN across Peak Count, Period Time, and Spike Spacing interpretations.
Figure 6. Prediction error distributions ( y^−y ) for each architecture–interpretation pair (rows: Peak Count, Period Time, Spike Spacing; columns: BiGRU, GRUReLU, TCN). A 3 by 3 grid of histograms showing the distribution of prediction errors for each combination of speed interpretation (rows) and architecture (columns).
For seemless control of advanced hand prostheses and augmented reality, accurate and immediate hand gestures recognition is essential. Surface electromyography (sEMG) signals obtained from the forearm are commonly employed for this purpose. In this paper, we present a novel approach for sEMG representation that utilizes graph networks which contain information about muscle activation patterns in the forearm. Based on these graph networks, we have developed a machine learning algorithm capable of real-time hand gesture recognition using a graph neural network. The algorithm's performance was evaluated using sEMG signals acquired from myoband, which has 8 electrodes placed around the forearm, involving 8 healthy subjects. The proposed method demonstrated an average classification accuracy of 99%, surpassing the performance of state-of-the-art techniques. The average time for both graph construction and prediction stood at 48ms utilizing a M1 pro CPU, rendering the approach well-suited for real-time applications.
Pragatheeswaran Vipulanandan, Kamal Premaratne, Manohar Murthi
Department of Electrical and Computer Engineering University of Miami Coral Gables, Florida, USA
In human computer interaction, real-time detection and classification of dynamic hand gestures is challenging as: 1) the system must run in a real-time video stream and there is no noticeable lag in response after performing a gesture; 2) there is a large difference in how people perform gestures, making recognition more difficult. In this paper, an online hand gesture recognition system is proposed, which is able to localize gestures in real-time video stream and recognize what these gestures are. To improve the robustness of the system, the sliding window approach is used to refine results from multiple windows. All of the models in my project are trained on Jester database, achieving 98+% accuracy for detector and 90+% accuracy for classifier. For the overall performance of the system, the best group can respond within three seconds and reach 37.5% Levenshtein accuracy on the homemade dataset. The project codes used in this work are publicly available.
Yinghao Qin, Tijana Timotijevic
School of Electronic Engineering and Computer Science Queen Mary, University of London London, UK
Continuous estimation of high-dimensional finger kinematics from forearm surface electromyography (EMG) could enable natural control for hand prostheses, AR/XR interfaces, and teleoperation. However, the complexity of human hand gestures and the entanglement of forearm muscles make accurate recognition intrinsically challenging. Existing approaches typically reduce task complexity by relying on classification-based machine learning, limiting the controllable degrees of freedom and compromising on natural interaction. We present an end-to-end framework for continuous EMG-to-kinematics regression using only consumer-grade hardware. The framework combines an 8-channel EMG armband, a single webcam, and an automatic synchronization procedure, enabling the collection of the EMG Finger-Kinematics dataset (EMG-FK), a 10-h dataset of synchronized EMG and 15 finger joint angles from 20 participants performing rich, unconstrained right-hand motions. We also introduce the Temporal Riemannian Regressor (TRR), a lightweight GRU-based model that uses sequences of multi-band Riemannian covariance features to decode finger motion. Across EMG-FK and the public emg2pose benchmark, TRR outperforms state-of-the-art methods in both intra- and cross-subject evaluation. On EMG-FK, it reaches an average absolute error of 9.79°±1.48 in intra-subject and 16.71°±3.97 in cross-subject. Finally, we demonstrate real-time deployment on a Raspberry Pi 5 and intuitive control of a robotic hand; TRR runs at nearly 10 predictions/s and is roughly an order of magnitude faster than state-of-the-art approaches. Together, these contributions lower the barrier to reproducible, real-time EMG-based decoding of high-dimensional finger motion, and pave the way toward more natural and intuitive control of embedded EMG-based systems.
Martin Colot, Cédric Simar, Guy Cheron +2
Machine Learning Group, ULB, Bruxelles, Belgium · § These authors contributed equally to this work and share first authorship. · Neuromove, ULB Neuroscience Institute, ULB, Brussels, Belgium +1