Deploying intelligent robotic systems that interact with humans through gestures requires neural networks capable of recognizing diverse temporal patterns. We present a systematic benchmark of ten abstract sequential tasks--five permutation-invariant (set) and five order-dependent (sequence) problems--evaluated across eighteen neural network architectures spanning recurrent, convolutional, attention-based, and set-function families. Beyond the core architecture-task grid, we explore numerous preprocessing and target-variable transformations, yielding more than 250 distinct experimental configurations. All variants are trained and tested under strictly identical conditions (fixed random seeds, shared hyperparameters, shared data splits) to ensure fair and reproducible comparison. Ranking across all ten tasks reveals four consistently top-performing architectures--BiGRU, TCN, Conv1D, and GRUReLU--all compact enough for real-time deployment (under 2,000 parameters in the benchmark setting). Based on this ranking, we apply three architecturally diverse top models (BiGRU, TCN, and GRUReLU) to a practical robotics problem: estimating the execution speed of dynamic arm gestures from skeletal keypoint sequences. Three speed interpretations (peak count, period time, and mean spike spacing) are evaluated on a custom dataset of eight traffic-related gesture classes comprising 256,710 frames recorded via OpenPose. The best configuration achieves a mean absolute error of 0.198 on the peak-count interpretation, corresponding to roughly 5% relative error, while the period-time interpretation reaches approximately 4% relative error, and the mean spike spacing interpretation approximately 8% relative error. These results demonstrate that neural networks can reliably estimate gesture speed from skeletal data, opening a path toward speed-aware gesture-controlled robotic systems.
Figures & tables
Architecture
Params
1-Decision
2-Counting
3-Variance
4-Mean
5-Max
GRU
14
1.000
3.428 11
0.004
0.001 2
0.006 4
GRUReLU
963
0.994 2
1.226 9
0.004
0.002 3
0.011 9
BiGRU
1,921
1.000
2.020 10
0.004
0.001 2
0.007 5
LSTM
16
0.960 5
12.726 14
0.025 4
0.012 7
0.020 11
BiLSTM
2,529
0.988 4
14.126 15
0.012 3
0.015 9
0.033 13
SkipGRU
14
1.000
19.444 16
0.062 9
0.014 8
0.044 14
Table 1. Performance on set tasks ( 1000:1000 configuration). Task 1: accuracy ( ↑ ); Tasks 2–5: MAE ( ↓ ). Values rounded to 3 decimal places. Bold = rank 1. Superscript = per-task dense rank (ties share the same rank). – = convergence failure (ranked last).
Architecture
Params
6-Speed
7-Peak
8-Period
9-SpikeSp
10-Threshold
GRU
14
0.004
225.496 18
2.500 13
6.752 2
10.318 13
GRUReLU
963
0.006 3
123.024 11
0.074 4
6.752 2
6.294 2
BiGRU
1,921
0.008 4
8.023 3
0.031 2
6.752 2
6.243
LSTM
16
0.083 9
225.267 17
2.500 13
6.753 3
133.807 16
BiLSTM
2,529
0.069 7
113.481 10
0.092 5
6.752 2
6.302 3
SkipGRU
14
0.063 6
180.412 15
2.501 14
6.759 6
241.764 17
Table 2. Performance on sequence tasks ( 1000:1000 configuration, MAE ↓ ). Values rounded to 3 decimal places. Bold = rank 1. Superscript = per-task dense rank (ties share the same rank). – = convergence failure (ranked last).
#
Model
Params
Σrank
ΣSET
ΣSEQ
1
BiGRU
1,921
31
19
12
1
TCN
1,667
31
19
12
2
Conv1D
292
43
20
23
3
GRUReLU
963
46
24
22
4
Transformer
24,961
54
19
35
5
Quantile
385
59
21
38
Table 3. All architectures ranked by global rank sum ( Σrank , lower is better). ΣSET and ΣSEQ show rank sums for set and sequence tasks, respectively.
Figure 1. Per-task dense rank of all 18 architectures across the 10 benchmark tasks ( 1000:1000 configuration). Each point shows one architecture’s rank on a single task; legend entries are ordered by global rank sum ( Σ ). Blue background: permutation-invariant set tasks (Tasks 1–5); orange: order-dependent sequence tasks (Tasks 6–10). Bump chart with one colored series per architecture, showing each model's per-task rank across the ten benchmark tasks.
Figure 2. Cross-task performance profiles of the five architectures with the lowest global rank sum ( 1000:1000 configuration). Spoke values are min–max normalized across all 18 architectures (1.0 = best, 0.0 = worst). Blue sector: set tasks; orange sector: sequence tasks. Radar chart showing ten task spokes; each of the five top-ranked architectures is drawn as a colored polygon indicating its normalized score on every task.
Figure 3. Reference poses for each of the eight gesture types, visualized as BODY-25 skeletal keypoints after 1×1 normalization. Blue: right arm; red: left arm; black: torso; green: head. Eight skeleton plots showing the normalized reference position for each gesture type used in the speed estimation application.
Figure 4. Weighted distance-from-reference-pose curves for example gesture sequences. Red dots mark local minima—moments at which the skeleton is closest to the reference pose. Inter-minima spacing reflects execution speed. Line plots showing the weighted distance from the reference pose over time for different gestures, with local minima marked by red dots.
Interpr.
Model
Params
MAE
MSE
Bias
Std.
7-Peak
BiGRU
27,649
0.198
0.097
+ 0.009
0.311
GRUReLU
13,825
0.214
0.113
− 0.009
0.336
TCN
6,785
0.272
0.141
+ 0.065
0.370
8-Period
GRUReLU
13,057
1.320
5.942
− 0.220
2.428
BiGRU
26,113
1.499
7.157
+ 0.478
2.632
TCN
6,401
2.882
16.756
+ 0.775
4.019
Table 4. Application results on real gesture speed estimation. Bold = best per interpretation group: MAE, MSE, Std ( ↓ ); Params ( ↓ fewest); Bias ( ∣⋅∣ closest to zero). Bias is the signed mean error; Std is the error standard deviation.
Figure 5. MAE and MSE for the three selected architectures across the three speed interpretations. Grouped bar chart comparing the MAE of BiGRU, GRUReLU, and TCN across Peak Count, Period Time, and Spike Spacing interpretations.
Figure 6. Prediction error distributions ( y^−y ) for each architecture–interpretation pair (rows: Peak Count, Period Time, Spike Spacing; columns: BiGRU, GRUReLU, TCN). A 3 by 3 grid of histograms showing the distribution of prediction errors for each combination of speed interpretation (rows) and architecture (columns).
Machine Learning Group, ULB, Bruxelles, Belgium · § These authors contributed equally to this work and share first authorship. · Neuromove, ULB Neuroscience Institute, ULB, Brussels, Belgium +1