cs.CVMar 13, 2026

HARMONI: Aligning Human and Scene Priors for Multi-View 4D Reconstruction

Authors: Sangmin Kim, Minhyuk Hwang, Geonho Cha, Dongyoon Wee, Jaesik Park

Organizations: Seoul National University · NAVER Cloud

Abstract

Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independently trained human and scene priors often produces misalignment in scale and depth. We observe that the two priors have complementary strengths. The scene prior provides consistent depth but approximate scale, while the human prior provides fixed body scale but less reliable depth. Motivated by this observation, we present HARMONI, a feed-forward framework that reconstructs cameras, scene, and humans with identities from monocular or multi-view video without test-time optimization. We introduce bidirectional anchoring, in which scene depth guides human placement, while human keypoints calibrate the scene scale to match the human body. While most existing approaches target monocular inputs and multi-view methods rely on optimization or re-identification, our approach naturally extends to multiple views. Building on bidirectional anchoring, we introduce a multi-view fusion module that merges per-view estimates, while reducing the influence of unreliable views and visually similar individuals. Experiments show that our framework outperforms previous human-scene methods in global motion and multi-view pose estimation by up to 28% and 64%, while running 28 times faster than optimization-based approaches. Project page: https://nstar1125.github.io/harmoni.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

    Jun 1, 2026Jinpeng Liu, Yukang Xu, Yutong Li +1Scene ReconstructionCamera

  2. MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images

    Jun 11, 2025Chentao Song, He Zhang, Haolei Yuan +4Human Mesh RecoveryScene Reconstruction

  3. Scene and Human in One World: Reconstruction in a Feedforward Pass

    Jun 26, 2026Boao Shi, Qiao Feng, Yiming Huang +1Human Mesh RecoveryMonocular Video