cs.CVSep 24, 2026

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

Authors: Jerrin Bright, John Zelek

Organizations: Vision and Image Processing Lab, University of Waterloo, Canada

Abstract

3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through lightweight adaptation. Exact metric labels are scarce, but unlabeled in-the-wild video is abundant. We use people in curated web video to initialise the solution: a posed metric body and 2D keypoints give an approximate, closed-form scale pseudo-label. These pseudo-labels pretrain a Scale Readout, which is then fine-tuned together with a lightweight adapter using exact metric supervision from standard real-video training splits. At inference the head predicts metric scale from foundation-model tokens, without the ruler or its teachers. For person identity, we probe the pretrained foundation model alone and find evidence that its intermediate query-key features encode person correspondence across frames. In most evaluated moving-person clips, a mid-layer token prefers that person over the vacated location and other people. A tiny projection reads this correspondence; together with metric pelvis motion and proposal confidence, it drives dustbin-aware Sinkhorn association of per-frame bodies. WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video. Each window is predicted feed-forward; analytic association and Sim(3) composition connect windows. On EMDB-2, WildHSR is the first feed-forward method in the published comparison to beat the best optimization-based WA-MPJPE and RTE while leading feed-forward methods on all three world-frame metrics. On RICH, it leads feed-forward people-and-scene methods on WA-MPJPE and W-MPJPE. The complete pipeline runs at 10.1 fps on one GPU.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scene and Human in One World: Reconstruction in a Feedforward Pass

    Jun 26, 2026Boao Shi, Qiao Feng, Yiming Huang +1Human Mesh RecoveryMonocular Video

  2. HARMONI: Aligning Human and Scene Priors for Multi-View 4D Reconstruction

    Mar 13, 2026Sangmin Kim, Minhyuk Hwang, Geonho Cha +2Multi-View4D Reconstruction

  3. MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images

    Jun 11, 2025Chentao Song, He Zhang, Haolei Yuan +4Human Mesh RecoveryScene Reconstruction