cs.CVOct 6, 2026

RenderBench: Benchmarking Render-to-Real Video Transfer with Reconstructed Digital Twins

Authors: Dicong Qiu, Zhiyuan Xu, Yaosheng Liu, Feng Han, Bo Ye

Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Southeast University · Independent Researcher

Abstract

Modern video models can generate realistic videos from real appearance references and proxy renders that specify scene structure, viewpoint changes, and motion. Evaluating this render-to-real capability requires a real target video depicting the same scene evolution, paired with an editable, geometrically registered 3D replica. Such data has traditionally required substantial manual modeling, calibration, and animation effort. We introduce RenderBench, a benchmark of 12 reconstructed real-world scenes spanning large-scale indoor environments and egocentric viewpoints, with both static and dynamic settings. Our construction pipeline combines visual geometry, neural reconstruction, and assisted 3D authoring. Each scene is decomposed into static objects and dynamic actors, registered to the capture cameras, and accepted only after multi-view geometric and temporal validation. Each evaluation unit contains appearance reference images, a held-out real target video, an editable digital twin, a matched proxy render, and renderer-native scene annotations. We evaluate transfer models against paired real target videos, retain PAI-Bench-C-compatible structural projections, and use scene annotations to localize failures by object, visibility, articulation, and motion. The first release retains 12 of 14 registered samples (85.7%), comprising 1,496 paired real-proxy frames. All released scenes pass file-integrity and environment-edit audits, while proxy diagnostics yield a depth si-RMSE of 0.2170 and instance mIoU of 0.3673. RenderBench provides paired real observations and editable scene state for assessing both appearance fidelity and preservation of geometry and dynamics.

Figures & tables

Explore similar work

CardsList
  1. GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction

    Jun 23, 2026Chenrui Fan, Paolo FavaroText-To-Video Generation ModelRobust Geometric Model Estimation

  2. Quantitative Video World Model Evaluation for Geometric-Consistency

    May 14, 2026Jiaxin Wu, Yihao Pi, Yinling Zhang +2Generative Video ModelsVideo World Models