cs.CVSep 30, 2026

PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

Authors: Rikhat Akizhanov, Yangsong Zhang, Nikolai Kaliazin, Peter Wolf, Yoshihiko Nakamura, Pascal Fua, Fabio Pizzati, Ivan Laptev

Organizations: MBZUAI · ETH Zürich · EPFL

Abstract

Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

    Apr 21, 2026Tanuj Sur, Shashank Tripathi, Nikos Athanasiou +4Scene ReconstructionMonocular Video

  2. Recovering Physically Plausible Human-Object Interactions from Monocular Videos

    Jun 3, 2026Dingbang Huang, Etienne Vouga, Qixing Huang +1Monocular VideoMotion