cs.CVOct 8, 2026

EgoPhys: Estimating Peak Contact Force and Mechanical Work from Egocentric Manipulation Video

Authors: Zhuo Dong, Jianhua Yang, Haohao Li, Yumeng Zhao, Keji He, Yan Huang, Liang Wang

Organizations: School of Artificial Intelligence, Shandong University, Jinan, China · Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Mechanical Engineering, Tianjin University, Tianjin, China

Abstract

Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because physical interaction cues are local and indirect. Moreover, peak force is associated with brief contact events, whereas mechanical work depends on force-motion coupling throughout the contact duration. To address these challenges, we propose EgoPhys, an RGB-only framework comprising Contact-Aware Spatial Aggregation (CASA) and Target-Specific Multi-Expert Temporal Routing (TMTR). CASA integrates appearance and geometry features to emphasize interaction-relevant cues, while TMTR models semantic, event, and motion cues with specialized temporal experts and routes them separately for force and work prediction. On the test split from Hoi! dataset, EgoPhys substantially improves predictions of peak force and mechanical work, achieving MAEs of 5.205±0.5845.205 \pm 0.584 NN and 0.894±0.0810.894 \pm 0.081 JJ, respectively.

Figures & tables

Explore similar work

CardsList
  1. EgoPHI: Estimating 3D Hand-Object Contact and Force from Egocentric Vision

    Aug 13, 2026Andela Ilic, Rachel Schuchert, Yijing Jiang +1Egocentric VisionHuman-Object Interaction

  2. PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video

    Sep 30, 2026Rikhat Akizhanov, Yangsong Zhang, Nikolai Kaliazin +5Human Pose Estimation3D Human Reconstruction

  3. EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video

    Jun 8, 2026Yuan Zeng, Yujia Shi, Tiao Tan +6Vision-Based Tactile SensingVideo Diffusion Models