EgoPhys: Estimating Peak Contact Force and Mechanical Work from Egocentric Manipulation Video
Authors: Zhuo Dong, Jianhua Yang, Haohao Li, Yumeng Zhao, Keji He, Yan Huang, Liang Wang
Organizations: School of Artificial Intelligence, Shandong University, Jinan, China · Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Mechanical Engineering, Tianjin University, Tianjin, China
Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because physical interaction cues are local and indirect. Moreover, peak force is associated with brief contact events, whereas mechanical work depends on force-motion coupling throughout the contact duration. To address these challenges, we propose EgoPhys, an RGB-only framework comprising Contact-Aware Spatial Aggregation (CASA) and Target-Specific Multi-Expert Temporal Routing (TMTR). CASA integrates appearance and geometry features to emphasize interaction-relevant cues, while TMTR models semantic, event, and motion cues with specialized temporal experts and routes them separately for force and work prediction. On the test split from Hoi! dataset, EgoPhys substantially improves predictions of peak force and mechanical work, achieving MAEs of 5.205±0.584N and 0.894±0.081J, respectively.
Figures & tables
Figure 1: Task overview. Egocentric RGB video is captured by a head-mounted camera, while force/torque and tool-motion signals are recorded by the Hoi! custom gripper. The task is to predict peak contact force and contact-phase mechanical work from the egocentric RGB video alone.
Figure 2: Overview of our proposed EgoPhys framework. Given an egocentric RGB clip, frozen appearance and geometry encoders extract complementary visual features. CASA aggregates interaction-relevant spatial evidence, while TMTR models complementary temporal cues to predict peak contact force and contact-phase mechanical work.
Peak Contact Force
Methods
MAE ( N ) ↓
MedianAE ( N ) ↓
RMSE ( N ) ↓
MRE ↓
DINOv3-only
9.320 ± 0.189
5.696 ± 0.337
13.827 ± 0.290
0.584 ± 0.015
DA3-only
9.318 ± 0.770
5.702 ± 0.523
13.689 ± 1.334
0.596 ± 0.039
DINOv3+DA3
8.796 ± 1.230
5.702 ± 0.801
12.689 ± 1.949
0.600 ± 0.068
Full Model
5.205 ± 0.584
3.492 ± 0.594
7.132 ± 0.765
0.301 ± 0.031
Mechanical Work
Table 1: Comparison with baseline methods on test split.
Figure 3: Qualitative results on the selected test video clip.
Peak Contact Force
Mechanical Work
Models
MAE ( N ) ↓
MRE ↓
MAE ( J ) ↓
MRE ↓
#A
8.796 ± 1.230
0.600 ± 0.068
1.624 ± 0.124
2.034 ± 0.337
#B
7.776 ± 0.308
0.447 ± 0.024
1.456 ± 0.210
2.062 ± 0.476
#C
5.982 ± 0.299
0.335 ± 0.030
1.103 ± 0.169
1.428 ± 0.222
Full model
5.205 ± 0.584
0.301 ± 0.031
0.894 ± 0.081
1.237 ± 0.189
Table 2: Ablation of the main components.
Peak Contact Force
Mechanical Work
Expert
MAE ( N ) ↓
MRE ↓
MAE ( J ) ↓
MRE ↓
Semantic
5.753 ± 0.267
0.319 ± 0.018
1.093 ± 0.074
1.467 ± 0.219
Event
6.401 ± 0.396
0.346 ± 0.036
1.344 ± 0.172
1.801 ± 0.269
Motion
7.039 ± 0.570
0.432 ± 0.043
1.571 ± 0.067
1.883 ± 0.314
All (Concat)
6.303 ± 0.263
0.359 ± 0.023
1.189 ± 0.037
1.578 ± 0.232
All (Gate)
5.205 ± 0.584
0.301 ± 0.031
0.894 ± 0.081
1.237 ± 0.189
Table 3: Comparison of temporal expert configurations.
Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet reasoning about physically grounded interaction requires estimating the forces acting on hands and objects, beyond localizing contact. We present EgoPHI, the first method that jointly estimates dense contact maps and 3D force distributions on hand and object meshes from a single monocular RGB image and object geometry. To address the lack of scalable ground-truth force annotations, we introduce a physics-based simulation pipeline that augments existing hand-object datasets with dense per-vertex force supervision. EgoPHI then learns dense 3D contact and force on interacting hand and articulated object meshes, extending vision-based force estimation beyond image-space or planar settings. Our evaluation on in-distribution and out-of-distribution benchmarks shows that EgoPHI improves force estimation over existing approaches while generalizing to unseen datasets. To evaluate sim-to-real transfer, we constructed two physical objects that capture dense object contact and force magnitude and used them to record a dataset of interactions from eight participants across diverse touch and grasp types. Our results demonstrate that EgoPHI recovers meaningful 3D contact and force distributions in simulated, out-of-distribution, and real-world settings, advancing egocentric hand-object understanding from contact localization toward physically grounded interaction reasoning.
Andela Ilic, Rachel Schuchert, Yijing Jiang +1
Department of Computer Science, ETH Zürich, Zürich, Switzerland
Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.
Rikhat Akizhanov, Yangsong Zhang, Nikolai Kaliazin +5
Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3D object interactions. Therefore, we introduce EgoTactile, a benchmark pairing egocentric video with full-hand pressure supervision for diverse everyday objects, incorporating a bare-hand transfer subset to enable generalization to natural scenarios. Leveraging this benchmark, we first establish EgoPressureFormer as a discriminative baseline. Beyond this, to explicitly address the uncertainty in partial observations, we propose EgoPressureDiff, a conditional diffusion framework that adapts a large-scale pre-trained video diffusion backbone. By combining rich world knowledge priors with a Physically-Informed Feature Rectification layer to inject semantic constraints, our approach effectively infers plausible contact patterns and resolves visual-physical ambiguities. Extensive experiments demonstrate that our method achieves superior performance on the benchmark and robust transferability to in-the-wild scenarios. Our project page is available at https://egotactile.github.io/.
Yuan Zeng, Yujia Shi, Tiao Tan +6
Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen, China · Department of Network, Pengcheng Laboratory, Shenzhen, China +2