HUMAN-TCI: Hierarchical Multi-Stream Motion-Aware Network with Torso-Centered Interaction for Text-to-Motion Retrieval
Authors: Muhammad Islam, Euijoon Ahn, Usman Naseem, Tao Huang
Organizations: College of Science and Engineering, James Cook University, Cairns, 4878, QLD, Australia · Centre for AI and Data Science Innovation, James Cook University, Cairns, 4878, QLD, Australia · School of Computing, Macquarie University, Sydney, 2113, NSW, Australia
Accurate retrieval of human motions is a crucial first step in text-guided human motion modeling and synthesis, as it selects semantically relevant sequences from large datasets and provides grounded references for downstream tasks. Retrieving motions from natural language descriptions remains challenging because sentences can describe multiple actions, overlapping movements, and intricate dependencies between body parts. Existing methods often focus on simple, single-action descriptions and typically process body parts independently or by merely concatenating features, without explicitly modeling how torso movements influence other parts. In addition, their processing pipelines often rely on computationally heavy models, introducing considerable overhead, particularly when modeling longer or more complex motion sequences. This limits learning discriminative motion-pattern representations, reducing retrieval accuracy, interpretability, and efficiency in practical applications. To address these limitations, we propose HUMAN-TCI, a Hierarchical Multi-Stream Motion-Aware Network for text-guided human motion retrieval. HUMAN-TCI employs a three-stream architecture that separately models upper-body, lower-body, and torso motions while explicitly capturing their interactions, allowing torso-related movements to influence the positioning and dynamics of other body parts. By incorporating tailored torso attention, our model effectively recognizes complex human motion patterns, captures fine-grained motion relationships and handles complex multi-action descriptions. Our framework supports retrieval for both simple, single-action sentences and long, compositional descriptions containing sequential or overlapping actions without relying on complex models.
Figures & tables
Figure 1: Overview of the proposed 3Tier-Hierarchical Architecture for text-to-motion retrieval. The framework learns hierarchical body-part-aware representations through dedicated upper-body, lower-body, and torso motion streams. Cross-stream dependencies are captured through torso-aware interaction (Figure 3), enabling fine-grained text–motion alignment in a shared embedding space. Triangles and circles represent text (T) and motion (M) embeddings, respectively, in the higher-dimensional embedding space. Details of the motion encoder are shown in Figure 2, while the text encoder is presented in Figure 4.
Figure 2: Illustration of the HUMAN-TCI motion pipeline. The input 3D skeletal sequence is first preprocessed and reduced from all joints (21-kitml, 22 H-ml3d) to 5 anatomically meaningful body regions, which are then organized into three body-part streams: upper body, torso, and lower body. The three streams are projected into a shared feature space and passed to the Torso-Center Interaction (TCI) module, detailed in Figure 3, where torso-aware attention models cross-stream dependencies between the upper-body, torso, and lower-body representations. The resulting attended features are then independently encoded using GRU layers to capture temporal motion patterns. Finally, the encoded upper-body, torso, and lower-body features are concatenated and passed through a nonlinear projection layer followed by L2 normalization to obtain the final motion embedding for text-to-motion retrieval.
Figure 3: HUMAN-TCI module. Motion is decomposed into upper, lower, and torso streams. Upper/lower streams form queries, while the torso provides keys/values for attention. The attended features are encoded with GRU layers and fused into a unified motion embedding.
Figure 4: Detail proposed text encoding pipeline in HUMAN-TCI. A natural language description is first tokenized and encoded using BERT to obtain contextualized token-level hidden representations. The resulting sequence of BERT hidden states is then passed to a bidirectional LSTM to model sequential dependencies and capture long-range contextual information. The resulting text representation is subsequently projected into the shared motion–text embedding space, where it is aligned with the corresponding motion representation for fine-grained text-to-motion retrieval.
Method
Pub. Year
KIT-ML
HumanML3D
R@1 ↑
R@5 ↑
R@10 ↑
MedR ↓
R@1 ↑
R@5 ↑
R@10 ↑
MedR ↓
T2M [ 6 ]
CVPR’22
3.37
16.87
27.71
28
1.80
7.12
12.47
81
MotionCLIP [ 20 ]
ECCV’22
4.87
20.09
31.57
26
2.33
12.77
18.14
103
TEMOS [ 15 ]
ECCV’22
7.11
24.10
35.66
24
2.12
8.26
13.52
173
MoT [ 12 ]
SIGIR’23
6.23
23.92
37.15
20
2.61
10.66
17.79
60
TMR [ 16 ]
ICCV’23
7.23
28.31
40.12
17
5.68
20.34
30.94
28
Table 1: Comparison with state-of-the-art methods on KIT-ML [ 17 ] and HumanML3D [ 6 ] .
Motion Model
R@1 ↑
R@5 ↑
R@10 ↑
MedR ↓
MeanR ↓
KIT
Hier-2TGRU
3.71
15.22
28.57
31.00
72.34
Hier-3TGRU
5.44
21.83
36.20
21.00
61.54
Hier-3TGRU-Att
7.86
29.65
44.48
14.00
26.97
Hier-3TGRU-Att-HNP
9.96
32.31
47.07
13.00
18.51
HumanML3D
Table 2: Ablation study on KIT and HumanML3D datasets.
Text Model
KIT
HumanML3D
R@1 ↑
R@5 ↑
R@10 ↑
R@1 ↑
R@5 ↑
R@10 ↑
BERT-Large
7.82
29.57
44.27
6.84
25.1
36.54
CLIP
9.96
32.31
47.08
8.21
27.17
38.87
Table 3: Ablation study on KIT and HumanML3D datasets: BERT-LSTM vs CLIP text models. Best results in bold.
Model
Body Parts ( J )
Batch Size
Time (m)
KIT-ML
HumanML3D
BERT-LSTM + GRU
5
32
4.47
4.82
CLIP + GRU
5
32
2.36
2.51
Full Joint Model
22
32
6.12
6.45
Table 4: Computational efficiency comparison of text-to-motion retrieval models on KIT-ML and HumanML3D. Full inference time is reported in minutes.
Figure 5: 3D skeleton visualization depicting the spatiotemporal evolution of human motion across sequential frames. The visualization represents the temporal progression of articulated body joints, highlighting variations in body posture, joint movements, and structural coordination during motion execution. Such skeletal representations provide a compact view of human dynamics for learning discriminative motion features.
Figure 6: Visualization of human motion using the SMPL body model, showing the reconstructed 3D human mesh over sequential frames. The SMPL-based representation preserves detailed body geometry and pose variations while capturing the spatiotemporal dynamics of articulated human movements, enabling an intuitive interpretation of learned motion patterns.
Figure 7: Detailed torso-centered attention analysis for the query “A person takes five slow forward steps.” Top row: per-limb attention distribution over torso frames for Right Arm, Left Arm, Right Leg, and Left Leg, with peak attention frame ( t ) and Shannon entropy ( H ) annotated for each. Middle-left: attention heatmap across all four limbs (rows) and torso frames (columns), row-normalized for visual contrast, with diamond markers indicating each limb’s peak-attention frame. Middle-right: peak (solid) versus mean ± std (light) attention magnitude per limb. Bottom-left: all four limb attention curves overlaid for direct comparison. Bottom-right: total attention (summed across limbs) and its cumulative distribution over the sequence, with F50 and F90 marking the frames by which 50% and 90% of cumulative attention mass is reached, respectively.
Model
Dataset
nDCG ↑
SPICE
spaCy
Upper-Lower GRU
KIT-ML
0.271
0.706
Hier-3TGRU
KIT-ML
0.263
0.697
Hier-3TGRU + CLIP
KIT-ML
0.339
0.770
Upper-Lower GRU
HumanML3D
0.318
0.723
Hier-3TGRU
HumanML3D
0.316
0.729
Table 5: Semantic evaluation of text-to-motion retrieval models on KIT-ML and HumanML3D. Higher values indicate better semantic alignment.
Figure 8: Visualization of Torso-Centered Interaction learned by our proposed AttentionFuse module. In our 3T-hierarchical architecture, each limb branch (arms, legs) queries the torso branch to selectively retrieve temporally-relevant torso context, rather than encoding limb motion in isolation. For nine representative test queries, we plot the resulting torso-attention distribution: the x-axis is the torso frame index and the y-axis is the normalized attention weight (%) each limb assigns to that frame. The bold navy curve is the mean attention across all four limbs (aggregate torso reliance); the thin red/green curves show per-limb (arm/leg) attention. Orange diamonds mark the frames of peak torso reliance, automatically extracted via cumulative-attention sampling.
Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search across the two modalities. Most existing methods use a dual-encoder framework that compresses motion and text into global embeddings, discarding fine-grained local correspondences, and thus reducing accuracy. Additionally, these global-embedding methods offer limited interpretability of the retrieval results. To overcome these limitations, we propose an interpretable, joint-angle-based motion representation that maps joint-level local features into a structured pseudo-image, compatible with pre-trained Vision Transformers. For text-to-motion retrieval, we employ MaxSim, a token-wise late interaction mechanism, and enhance it with Masked Language Modeling regularization to foster robust, interpretable text-motion alignment. Extensive experiments on HumanML3D and KIT-ML show that our method outperforms state-of-the-art text-motion retrieval approaches while offering interpretable fine-grained correspondences between text and motion. The code is available in the supplementary material.
Yao Zhang, Zhuchenyang Liu, Yanlan He +2
Aalto University, Espoo, Finland · Fudan University, Shanghai, China · Georgia Institute of Technology, Atlanta GA, USA
The world knowledge and reasoning capabilities of text-based large language models (LLMs) are advancing rapidly, yet current approaches to human motion understanding, including motion question answering and captioning, have not fully exploited these capabilities. Existing LLM-based methods typically learn motion-language alignment through dedicated encoders that project motion features into the LLM's embedding space, remaining constrained by cross-modal representation and alignment. Inspired by biomechanical analysis, where joint angles and body-part kinematics have long served as a precise descriptive language for human movement, we propose \textbf{Structured Motion Description (SMD)}, a rule-based, deterministic approach that converts joint position sequences into structured natural language descriptions of joint angles, body part movements, and global trajectory. By representing motion as text, SMD enables LLMs to apply their pretrained knowledge of body parts, spatial directions, and movement semantics directly to motion reasoning, without requiring learned encoders or alignment modules. We show that this approach goes beyond state-of-the-art results on both motion question answering (66.7% on BABEL-QA, 90.1% on HuMMan-QA) and motion captioning (R@1 of 0.584, CIDEr of 53.16 on HumanML3D), surpassing all prior methods. SMD additionally offers practical benefits: the same text input works across different LLMs with only lightweight LoRA adaptation (validated on 8 LLMs from 6 model families), and its human-readable representation enables interpretable attention analysis over motion descriptions. Code, data, and pretrained LoRA adapters are available at https://yaozhang182.github.io/motion-smd/.
Yao Zhang, Zhuchenyang Liu, Thomas Ploetz +1
Aalto University, Espoo, Finland · Georgia Institute of Technology, Atlanta, GA, USA
Text-to-motion generation has progressed rapidly in recent years, offering an expressive interface for animation and human-computer interaction. However, current models remain brittle when handling prompts that describe multiple actions occurring at the same time. Rather than realizing all components of a composite description, models frequently prioritize a single dominant action and neglect the rest, leading to incomplete or ambiguous motion. We present MultiAct, an unpaired, inference-time framework for compositional text-to-motion synthesis that operates directly on pretrained motion generators without retraining or architectural modification. Our method counteracts semantic collapse by adaptively amplifying cross-attention scores associated with underrepresented prompt components. We note that effective modulation depends on prompt-specific choices, such as which tokens and layers to target, and introduce a lightweight auxiliary decision scheme that determines the most effective attention-strengthening parametrization. Extensive quantitative and qualitative evaluations demonstrate that MultiAct consistently outperforms existing baselines on composite prompts, achieving improved semantic coverage while preserving motion realism. Project page: https://natsala13.github.io/multiact.github.io.
Nathan Sala, Ofir Abramovich, Ariel Shamir +3
Tel Aviv University, Israel · Reichman University, Israel · CYENS Centre of Excellence, Cyprus +1