This paper presents a video-to-model framework for automatically modeling the motion of a deformable suture thread from an input video. We utilize a recently developed CBF--CLF--QP numerical model that simplifies the characterization of deformable string motion through the selection of a small number of parameters. A perception module first localizes and tracks the thread in video, producing an ordered sequence of thread nodes. The observed thread motion is then processed by a spatio-temporal CNN network that estimates the effective parameters of a structured CBF--CLF--QP model. These parameters are used to simulate the thread under a user-defined needle velocity input. Experiments using unseen thread configurations and motion demonstrate that the framework can reliably reconstruct the thread behavior from video, automatically configure the structured model, and reproduce the expected thread motion with low tracking error. The proposed approach reduces the need for manual parameter tuning and provides a step toward automatic video-based modeling of deformable linear objects.
Figures & tables
Figure 1: Overview of the proposed video-to-model framework for automatic thread-model configuration. An input video is fed into a perception module to detect the centerline of the DLO. This object is then fed into a parameter estimator module which enables a CBF–CLF–QP simulation of the object.
Figure 2: Thread-localization and tracking pipeline. An input RGB frame is processed by the U-Net to obtain a thread-probability map, which is thresholded and cleaned using morphological closing and small-component removal. The resulting mask is skeletonized, pruned, and traced to obtain an ordered thread centerline. The green and red markers indicate the two centerline endpoints.
Figure 3: Different needle trajectories (thread tip) used to generate the training dataset.
Material
Estimate
γ
α
κ
QP error (mm)
1
Ground Truth
11.174
13.561
0.642
–
No Segmentation
10.460
15.633
0.507
0.060
U-Net
2.489
2.267
0.526
0.639
2
Ground Truth
9.896
1.058
0.623
–
No Segmentation
15.757
1.815
0.433
0.043
U-Net
1.105
0.832
0.681
0.095
Table 1: Parameter identification and QP replay for the asymmetric flower-petal trajectory. Errors are mean node-trajectory errors.
Figure 4: Asymmetric-flower-petal needle trajectory used for evaluation, comprising a forward flower-petal sweep, a variable-speed reverse, and a return to the initial position.
Figure 5: Representative thread configurations during the asymmetric flower-petal motion experiment. Black denotes the ground-truth trajectory, cyan denotes the U-Net observations, and magenta denotes the trajectory predicted by our model.
World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challenging because deformable linear, planar, and volumetric objects evolve under high-dimensional deformation, noisy interactions, and complex material response. The model must therefore infer a physical state from visual observations, roll it forward under new interactions, and render the resulting dynamics with high visual fidelity. We present DeformMaster, a video-derived interactive physics-neural world model that turns real interaction videos into an online interactive model of deformable objects within a unified dynamics-and-appearance framework. DeformMaster preserves structured physical rollout while using a neural residual to compensate for unmodeled effects, grounds sparse hand motion as distributed compliant actuator for hand-continuum interaction, represents material response with spatially varying constitutive experts, and drives high-fidelity 4D appearance from the predicted physical evolution. Experiments on real-world deformable-object sequences demonstrate DeformMaster's ability to roll out future dynamics and render dynamic appearance, outperforming state-of-the-art baselines while supporting novel action rollout, material-parameter variation, and dynamic novel-view synthesis. Project page: https://can-lee.github.io/deformmaster-web/
Can Li, Zhoujian Li, Ren Li +4
Nankai University · Rightly Robotics, A4X · Zhejiang University +2
Reconstructing objects with mechanical properties from video observations enables physically consistent dynamic prediction, benefiting robotics planning and interaction. Existing spring--mass based physical driven reconstruction approaches offer efficient and differentiable physical reconstruction, but they typically rely on axial springs alone. Such formulations oversimplify the underlying structural mechanics and can become mechanically under-constrained when the physical graph is coarsened, limiting their ability to preserve stable local deformation. We present BendTwin, a bending-aware differentiable spring--mass framework for video-based reconstruction and future prediction of deformable objects. BendTwin introduces bending stiffness and damping over local surface triplets, penalizing deviations from rest angles and regularizing higher-order deformation. These bending constraints improve mechanical stability while preserving the simplicity of spring--mass system. Experiments show that BendTwin consistently outperforms the axial-only PhysTwin baseline. Ablation studies further demonstrate that the bending constraints maintain system stability across different downsampling ratios and consistently improve upon the original PhysTwin formulation. Overall, BendTwin provides an effective approach for constructing mechanically faithful digital twins from sparse-view RGB-D videos.
Yixiong Jing, Qi Wang, Lin Chen +6
University of Cambridge · Institute of Automation, Chinese Academy of Sciences · Northwestern Polytechnical University +1
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.
Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu +6