HUMAN-TCI: Hierarchical Multi-Stream Motion-Aware Network with Torso-Centered Interaction for Text-to-Motion Retrieval
Authors: Muhammad Islam, Euijoon Ahn, Usman Naseem, Tao Huang
Organizations: College of Science and Engineering, James Cook University, Cairns, 4878, QLD, Australia · Centre for AI and Data Science Innovation, James Cook University, Cairns, 4878, QLD, Australia · School of Computing, Macquarie University, Sydney, 2113, NSW, Australia
Accurate retrieval of human motions is a crucial first step in text-guided human motion modeling and synthesis, as it selects semantically relevant sequences from large datasets and provides grounded references for downstream tasks. Retrieving motions from natural language descriptions remains challenging because sentences can describe multiple actions, overlapping movements, and intricate dependencies between body parts. Existing methods often focus on simple, single-action descriptions and typically process body parts independently or by merely concatenating features, without explicitly modeling how torso movements influence other parts. In addition, their processing pipelines often rely on computationally heavy models, introducing considerable overhead, particularly when modeling longer or more complex motion sequences. This limits learning discriminative motion-pattern representations, reducing retrieval accuracy, interpretability, and efficiency in practical applications. To address these limitations, we propose HUMAN-TCI, a Hierarchical Multi-Stream Motion-Aware Network for text-guided human motion retrieval. HUMAN-TCI employs a three-stream architecture that separately models upper-body, lower-body, and torso motions while explicitly capturing their interactions, allowing torso-related movements to influence the positioning and dynamics of other body parts. By incorporating tailored torso attention, our model effectively recognizes complex human motion patterns, captures fine-grained motion relationships and handles complex multi-action descriptions. Our framework supports retrieval for both simple, single-action sentences and long, compositional descriptions containing sequential or overlapping actions without relying on complex models.
Figures & tables
Figure 1: Overview of the proposed 3Tier-Hierarchical Architecture for text-to-motion retrieval. The framework learns hierarchical body-part-aware representations through dedicated upper-body, lower-body, and torso motion streams. Cross-stream dependencies are captured through torso-aware interaction (Figure 3), enabling fine-grained text–motion alignment in a shared embedding space. Triangles and circles represent text (T) and motion (M) embeddings, respectively, in the higher-dimensional embedding space. Details of the motion encoder are shown in Figure 2, while the text encoder is presented in Figure 4.
Figure 2: Illustration of the HUMAN-TCI motion pipeline. The input 3D skeletal sequence is first preprocessed and reduced from all joints (21-kitml, 22 H-ml3d) to 5 anatomically meaningful body regions, which are then organized into three body-part streams: upper body, torso, and lower body. The three streams are projected into a shared feature space and passed to the Torso-Center Interaction (TCI) module, detailed in Figure 3, where torso-aware attention models cross-stream dependencies between the upper-body, torso, and lower-body representations. The resulting attended features are then independently encoded using GRU layers to capture temporal motion patterns. Finally, the encoded upper-body, torso, and lower-body features are concatenated and passed through a nonlinear projection layer followed by L2 normalization to obtain the final motion embedding for text-to-motion retrieval.
Figure 3: HUMAN-TCI module. Motion is decomposed into upper, lower, and torso streams. Upper/lower streams form queries, while the torso provides keys/values for attention. The attended features are encoded with GRU layers and fused into a unified motion embedding.
Figure 4: Detail proposed text encoding pipeline in HUMAN-TCI. A natural language description is first tokenized and encoded using BERT to obtain contextualized token-level hidden representations. The resulting sequence of BERT hidden states is then passed to a bidirectional LSTM to model sequential dependencies and capture long-range contextual information. The resulting text representation is subsequently projected into the shared motion–text embedding space, where it is aligned with the corresponding motion representation for fine-grained text-to-motion retrieval.
Method
Pub. Year
KIT-ML
HumanML3D
R@1 ↑
R@5 ↑
R@10 ↑
MedR ↓
R@1 ↑
R@5 ↑
R@10 ↑
MedR ↓
T2M [ 6 ]
CVPR’22
3.37
16.87
27.71
28
1.80
7.12
12.47
81
MotionCLIP [ 20 ]
ECCV’22
4.87
20.09
31.57
26
2.33
12.77
18.14
103
TEMOS [ 15 ]
ECCV’22
7.11
24.10
35.66
24
2.12
8.26
13.52
173
MoT [ 12 ]
SIGIR’23
6.23
23.92
37.15
20
2.61
10.66
17.79
60
TMR [ 16 ]
ICCV’23
7.23
28.31
40.12
17
5.68
20.34
30.94
28
Table 1: Comparison with state-of-the-art methods on KIT-ML [ 17 ] and HumanML3D [ 6 ] .
Motion Model
R@1 ↑
R@5 ↑
R@10 ↑
MedR ↓
MeanR ↓
KIT
Hier-2TGRU
3.71
15.22
28.57
31.00
72.34
Hier-3TGRU
5.44
21.83
36.20
21.00
61.54
Hier-3TGRU-Att
7.86
29.65
44.48
14.00
26.97
Hier-3TGRU-Att-HNP
9.96
32.31
47.07
13.00
18.51
HumanML3D
Table 2: Ablation study on KIT and HumanML3D datasets.
Text Model
KIT
HumanML3D
R@1 ↑
R@5 ↑
R@10 ↑
R@1 ↑
R@5 ↑
R@10 ↑
BERT-Large
7.82
29.57
44.27
6.84
25.1
36.54
CLIP
9.96
32.31
47.08
8.21
27.17
38.87
Table 3: Ablation study on KIT and HumanML3D datasets: BERT-LSTM vs CLIP text models. Best results in bold.
Model
Body Parts ( J )
Batch Size
Time (m)
KIT-ML
HumanML3D
BERT-LSTM + GRU
5
32
4.47
4.82
CLIP + GRU
5
32
2.36
2.51
Full Joint Model
22
32
6.12
6.45
Table 4: Computational efficiency comparison of text-to-motion retrieval models on KIT-ML and HumanML3D. Full inference time is reported in minutes.
Figure 5: 3D skeleton visualization depicting the spatiotemporal evolution of human motion across sequential frames. The visualization represents the temporal progression of articulated body joints, highlighting variations in body posture, joint movements, and structural coordination during motion execution. Such skeletal representations provide a compact view of human dynamics for learning discriminative motion features.
Figure 6: Visualization of human motion using the SMPL body model, showing the reconstructed 3D human mesh over sequential frames. The SMPL-based representation preserves detailed body geometry and pose variations while capturing the spatiotemporal dynamics of articulated human movements, enabling an intuitive interpretation of learned motion patterns.
Figure 7: Detailed torso-centered attention analysis for the query “A person takes five slow forward steps.” Top row: per-limb attention distribution over torso frames for Right Arm, Left Arm, Right Leg, and Left Leg, with peak attention frame ( t ) and Shannon entropy ( H ) annotated for each. Middle-left: attention heatmap across all four limbs (rows) and torso frames (columns), row-normalized for visual contrast, with diamond markers indicating each limb’s peak-attention frame. Middle-right: peak (solid) versus mean ± std (light) attention magnitude per limb. Bottom-left: all four limb attention curves overlaid for direct comparison. Bottom-right: total attention (summed across limbs) and its cumulative distribution over the sequence, with F50 and F90 marking the frames by which 50% and 90% of cumulative attention mass is reached, respectively.
Model
Dataset
nDCG ↑
SPICE
spaCy
Upper-Lower GRU
KIT-ML
0.271
0.706
Hier-3TGRU
KIT-ML
0.263
0.697
Hier-3TGRU + CLIP
KIT-ML
0.339
0.770
Upper-Lower GRU
HumanML3D
0.318
0.723
Hier-3TGRU
HumanML3D
0.316
0.729
Table 5: Semantic evaluation of text-to-motion retrieval models on KIT-ML and HumanML3D. Higher values indicate better semantic alignment.
Figure 8: Visualization of Torso-Centered Interaction learned by our proposed AttentionFuse module. In our 3T-hierarchical architecture, each limb branch (arms, legs) queries the torso branch to selectively retrieve temporally-relevant torso context, rather than encoding limb motion in isolation. For nine representative test queries, we plot the resulting torso-attention distribution: the x-axis is the torso frame index and the y-axis is the normalized attention weight (%) each limb assigns to that frame. The bold navy curve is the mean attention across all four limbs (aggregate torso reliance); the thin red/green curves show per-limb (arm/leg) attention. Orange diamonds mark the frames of peak torso reliance, automatically extracted via cumulative-attention sampling.