cs.CVSep 28, 2026

HUMAN-TCI: Hierarchical Multi-Stream Motion-Aware Network with Torso-Centered Interaction for Text-to-Motion Retrieval

Authors: Muhammad Islam, Euijoon Ahn, Usman Naseem, Tao Huang

Organizations: College of Science and Engineering, James Cook University, Cairns, 4878, QLD, Australia · Centre for AI and Data Science Innovation, James Cook University, Cairns, 4878, QLD, Australia · School of Computing, Macquarie University, Sydney, 2113, NSW, Australia

Abstract

Accurate retrieval of human motions is a crucial first step in text-guided human motion modeling and synthesis, as it selects semantically relevant sequences from large datasets and provides grounded references for downstream tasks. Retrieving motions from natural language descriptions remains challenging because sentences can describe multiple actions, overlapping movements, and intricate dependencies between body parts. Existing methods often focus on simple, single-action descriptions and typically process body parts independently or by merely concatenating features, without explicitly modeling how torso movements influence other parts. In addition, their processing pipelines often rely on computationally heavy models, introducing considerable overhead, particularly when modeling longer or more complex motion sequences. This limits learning discriminative motion-pattern representations, reducing retrieval accuracy, interpretability, and efficiency in practical applications. To address these limitations, we propose HUMAN-TCI, a Hierarchical Multi-Stream Motion-Aware Network for text-guided human motion retrieval. HUMAN-TCI employs a three-stream architecture that separately models upper-body, lower-body, and torso motions while explicitly capturing their interactions, allowing torso-related movements to influence the positioning and dynamics of other body parts. By incorporating tailored torso attention, our model effectively recognizes complex human motion patterns, captures fine-grained motion relationships and handles complex multi-action descriptions. Our framework supports retrieval for both simple, single-action sentences and long, compositional descriptions containing sequential or overlapping actions without relying on complex models.

Figures & tables

Explore similar work

CardsList
  1. Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction

    Mar 10, 2026Yao Zhang, Zhuchenyang Liu, Yanlan He +2Text-To-Motion GenerationLatent Motion

  2. Encoder-Free Human Motion Understanding via Structured Motion Descriptions

    Apr 23, 2026Yao Zhang, Zhuchenyang Liu, Thomas Ploetz +1Human MotionHumanml3D

  3. MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance

    May 29, 2026Nathan Sala, Ofir Abramovich, Ariel Shamir +3Text-To-Motion GenerationCross-Attention