cs.CVSep 23, 2026

Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints

Authors: Mehmet Kerem TurkcanSoham SamalZoran Kostic

Abstract

Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument position, orientation and jaw angle from monocular video. Our visual representation combines global attention pooling of frozen DINOv3 features with local pooling at instrument landmarks from fine-tuned SAM 3.1 masks. Our shared Transformer encoder and temporal convolutional heads integrate this representation with mask geometry, monocular depth and visual state estimates from arm-specific multilayer regression networks. Our position branch predicts displacement magnitude and direction separately to preserve traveled distance. We fit trajectories to predicted state observations and motion increments by differentiable weighted least squares, expressing quaternion observations relative to cumulative predicted rotations to obtain a quadratic orientation objective. We evaluate reconstruction across 2,802 Open-H episodes. Compared with LiveMAE on the main Open-H benchmark, our method reduces path-length mean absolute error from 0.45 to 0.34,cm and increases temporal mean average precision for motion segmentation from 44.54% to 54.44%.

Explore similar work

CardsList
  1. MonoMSK: Monocular 3D Musculoskeletal Dynamics Estimation

    Nov 24, 2025Farnoosh Koleini, Hongfei Xue, Ahmed Helmy +1BiomechanicsInverse Kinematics