cs.CVSep 3, 2026

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

Authors: Yijun YangShenghe ZhengWenbo LiJianhui LiuHaoze SunYanbing ZhangJiaxiu JiangLin Song+3 more

Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Joy Future Academy · The Hong Kong University of Science and Technology · The University of Hong Kong

Abstract

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence (XYXY), depth consistency (ZZ), and temporal reversibility (TT). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.

Explore similar work

CardsList