cs.CVSep 24, 2026

PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark

Authors: Haodong Wang, Hezhen Hu, Wengang Zhou, Houqiang Li

Organizations: University of Science and Technology of China · University of Texas at Austin

Abstract

In this work, we focus on photorealistic sign avatar modeling, which is crucial for effective communication with the Deaf community and is characterized by complex hand gestures and nuanced facial expressions. To this end, we introduce MVSign, the first multi-view Chinese sign language dataset co-designed with Deaf experts, featuring diverse gestures and rich annotations. For precise SMPL-X annotation, we develop a hybrid fitting pipeline that produces accurate body, hand, and facial parameters and can also be applied to the monocular setting. Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity. Extensive experiments demonstrate that our method achieves high-fidelity visual results on MVSign, particularly in detailed hand and facial regions, and generalizes well to in-the-wild monocular sign language videos. Project page: https://naaapi.github.io/PHOSA.

Figures & tables

Explore similar work

CardsList
  1. Tamaththul3D: High-Fidelity 3D Saudi Sign Language Avatars from Monocular Video

    May 6, 2026Eyad Alghamdi, Sattam Altuuaim, Obay Ghulam +2Sign Language Translation3D Generation

  2. M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

    Mar 24, 2026Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan +1Sign Language TranslationExpression

  3. Towards Continuous Sign Language Conversation from Isolated Signs

    May 14, 2026Youngmin Kim, Kyobin Choo, Jiwoo Park +4Sign Language TranslationConversational Datasets