Proxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing Them
Organizations: Tencent
Abstract
Lightweight scene proxies let creators control scene layout and motion while leaving room for imagination in appearance, lighting, and visual effects. However, a suitable proxy is not uniquely defined, making paired proxy-video data difficult to construct automatically at scale. We present Proxy2World, a controllable world model that learns these complementary capabilities from ordinary posed RGBD videos, without training on authored proxy-video pairs. The model jointly learns depth-conditioned RGB generation and joint RGBD generation through cross-modal flow matching. Learning both tasks enables proxy-camera hybrid denoising at inference to follow the proxy structure while producing natural, detailed visuals. We further introduce ProxyBench to evaluate this capability across a diverse set of scenes, camera trajectories, and subject motions. Experiments on ProxyBench show that Proxy2World achieves a better balance between structural adherence and visual quality than camera-controlled and geometry-conditioned methods, supported by quantitative metrics, VLM assessments, human evaluations and diverse qualitative results.
Figures & tables
| Camera | Spatial Cons. | Proxy Evaluation | |||||||
| Operator-based | VLM-based | Human | |||||||
| Method | Input | Reproj. | Geo. Align | Temp. Align | Adh. WR | Ref. WR | Rank | ||
| Camera-conditioned methods | |||||||||
| Lyra 2.0 ( Shen et al., 2026 ) | T | 0.0418 | 0.0413 | 2.171 | 0.4965 | 0.7068 | 32.21 | 39.48 | 10 |
| SCoPE / RAYPE ( Yin et al., 2026 ) | T | 0.1066 | 0.1229 | 4.338 | 0.4508 | 0.5780 | 29.51 | 47.06 | 9 |
| FantasyWorld ( Dai et al., 2026 ) | T | 0.2966 | 0.3472 | 2.449 | 0.4711 | 0.6158 | 36.51 | 50.25 | 7 |