cs.CVSep 29, 2026

ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals

Authors: Xuyi Hu, Francesco Palandra, Shangzhe Wu, Daniel Cremers, Riccardo Marin, Silvia Zuffi

Organizations: University of Cambridge · IMATI-CNR, Milan, Italy · Technical University of Munich, Germany · Munich Center for Machine Learning, Germany

Abstract

Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual images and rely on synthetic or model-fitted 3D supervision, which inherits the constraints of strong parametric priors and limits generalization to out-of-distribution species. When applied to out-of-distribution animals, they often recover a plausible pose while producing inaccurate geometry because the underlying shape model cannot faithfully represent the observed instance. We present ORMA, a training-free reconstruction framework that decouples articulation from shape, using the predicted pose as reference for optimization while leveraging generative 3D priors for accurate shape reconstruction. Given a reference image, we reconstruct the animal geometry and register it to the parametric model SMAL+, yielding an articulated shape adapted to the observed instance. We then combine per-frame articulated pose estimates with globally consistent camera poses to recover animal motion in a shared world coordinate frame, and further refine the reconstruction using self-supervised DINO correspondences and temporal consistency. To enable quantitative evaluation, we introduce PAW4D, a synthetic multi-species benchmark with ground-truth 3D geometry and camera motion. Experiments on PAW4D, PFERD, and challenging in-the-wild videos demonstrate that ORMA improves reconstruction accuracy while recovering globally consistend animal motion across diverse quadruped species.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video

    Jun 30, 2026Siyuan Li, Weiying Chen, Yilin Wang +34D ReconstructionMonocular Video

  2. PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation

    Jun 1, 2026Xiaohang Yu, Ti Wang, Mackenzie Weygandt MathisHuman Mesh RecoverySkeleton

  3. One Video, One World: Turning Monocular Video into Physical 4D Scenes

    Jun 30, 2026Junhao Chen, Boran Zhang, Mingjin Chen +7Monocular Video4D