Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric errors. Our key insight is that geometry need not explain the scene--it only needs to determine where visual memory should be read from, while attention decides what should be recovered. Based on this insight, we introduce GEAR, a Geometry-Enabled Attention Routing framework that uses geometry as an explicit token-level address for visual memory. Rather than fusing historical observations into a persistent global 3D representation, GEAR retains them as frame latents and uses per-frame geometry only to establish token-level correspondences with target views, thereby avoiding persistent error accumulation from global fusion. Guided by these correspondences, Geometric Correspondence Attention (GCA) selectively injects geometrically matched historical features into noisy target patches during denoising. We further introduce an Invisible Octree to accumulate visibility evidence and reject geometrically plausible but occluded correspondences. Extensive experiments demonstrate that GEAR achieves state-of-the-art visual quality, precise camera control, and revisit consistency, enabling minute-long video generation along challenging trajectories.
Figures & tables
Figure 2 : Comparison of long-term memory paradigms. Implicit memory preserves rich visual history but requires dense search over an increasingly large token set, while explicit 3D memory provides spatial addressing at the cost of accumulated reconstruction and fusion errors. GEAR instead uses per-frame geometry only to address historical latent patches, enabling sparse, spatially grounded memory retrieval without persistent global 3D fusion.
Figure 3 : System overview. For each target chunk, GEAR constructs patch correspondences to retrieved history using per-frame geometry and filters occluded matches with the Invisible Octree. GCA then injects matched historical features into noisy target tokens during denoising, after which generated observations are appended to the history bank for continued rollout.
Figure 4 : Streaming update of the Invisible Octree. Historical observations incrementally accumulate coarse visibility evidence along the exploration trajectory. For a target view, the resulting invisible-space boundary is used to reject projectable but occluded source-target correspondences.
Figure 5 : Ablation of Invisible Octree. The Invisible Octree removes projectable but occluded correspondences, preventing erroneous historical content from affecting target-view synthesis.
Figure 6 : Qualitative comparison on DL3DV-Eval. Compared with explicit and implicit memory baselines, GEAR better preserves visual quality, camera adherence, and scene consistency throughout long-horizon generation. See our project page for additional scenes and video comparisons.
DL3DV-Evaluation
WorldScore-Static
Method
SSIM ↑
LPIPS ↓
FVD ↓
TransErr ↓
RotErr ↓
ATE ↓
Content Align. ↑
Photo. Cons. ↑
Style Cons. ↑
Subjective Quality ↑
Revisit SSIM ↑
Revisit LPIPS ↓
Lyra2
0.3359
0.5097
975.89
0.0157
0.1721
0.2514
0.6319
0.9357
0.8600
0.5017
0.3941
0.3218
Spatia
0.3081
0.5422
1074.13
0.0617
0.6973
1.1204
0.6331
0.8588
0.8600
0.5012
0.4407
0.3541
WorldStereo
0.3061
0.5502
846.46
0.0239
0.1717
0.2212
0.7003
0.0837
0.8300
0.5013
0.6253
0.2193
UCM
0.3412
0.6007
1431.67
0.0377
0.5184
0.6410
0.6773
0.9692
0.7300
0.5005
0.3412
0.4890
HY-WorldPlay
0.2452
0.6213
1385.23
0.0508
0.6854
0.9634
0.5104
0.7939
0.1900
0.5018
0.2433
0.7288
Table 1 : Quantitative comparison on DL3DV-Evaluation and WorldScore-Static. GEAR performs favourably across the metrics. Best results are in bold and second are underlined .
Figure 7 : Out-of-domain qualitative comparison for minute-long generation. GEAR maintains visual fidelity and scene consistency over extended camera trajectories, while competing methods exhibit progressive drift and visual degradation. *First-frame geometry projected along the target trajectory for viewpoint reference only. See our project page for additional video comparisons.
Figure 8 : Ablation of Degraded-History Augmentation. Training with degraded history mitigates error accumulation and improves visual stability during long-horizon autoregressive generation.
Figure 9 : Camera alignment. GEAR achieves the best camera alignment on DL3DV-Evaluation.
Figure 10 : Ablation of GCA. GCA improves temporal stability and adherence to the prescribed camera trajectory.
Figure 11 : Ablation of Correspondence Disturbance. We perturb candidates in multi-observed target regions by 64–128 pixels while retaining only 2–3 valid matches out of 9 conditions. The figure visualizes one such perturbation and the resulting error map; GCA recovers the target content with minimal deviation from the undisturbed result.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12 : Visualization of Geometry-Guided Patch Correspondence Construction. Local geometry from each historical observation is used to build a history-to-target patch correspondence cache at latent resolution. Guided by this cache, GCA allows each noisy target token to attend only to its geometrically matched historical memory tokens as keys and values.
Figure 13 : Visualization of the Invisible Octree update process. Partially visible voxels are recursively subdivided until each node becomes fully visible or fully invisible, or the maximum voxel resolution is reached. With increasing resolution, the Invisible Octree progressively approximates the true invisible regions in the target view.
Figure 14 : Visualization of the 3D pixel condition and generated frames of Spatia. Even with relatively clean 3D pixel-aligned conditioning, Spatia exhibits visible temporal instability and image degradation. Under more complex camera trajectories and noisier 3D pixel-aligned conditioning, frame jitter emerges in the first generated chunk, followed by a complete breakdown of visual content in the second.
Figure 15 : More results across a broader range of data.
Figure 16 : Qualitative comparison of minute-long challenging camera trajectories results
Figure 17 : Qualitative comparison of DL3DV-Evaluation results.
Figure 18 : Qualitative comparison of DL3DV-Evaluation results.
Figure 19 : Qualitative comparison of DL3DV-Evaluation results.
Figure 20 : Qualitative comparison of WorldScore-Static results.
Figure 21 : Qualitative comparison of WorldScore-Static results.
Figure 22 : Qualitative comparison of WorldScore-Static results.
Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models address this by retrieving historical frames, but their effectiveness depends on two key design choices: what 3D-geometric evidence should represent past observations, and how memory frames should be selected from this evidence. Existing methods often rely on camera poses or field-of-view overlap, which are lightweight but too coarse to reason about pixel-wise visibility, or use explicit 3D reconstruction, which provides fine-grained evidence but is costly to maintain over long rollouts. We propose Coverage-Maximizing Retrieval-Augmented Generation (COVRAG), a depth-based memory retrieval framework that uses pretrained 3D priors to construct a target-view coverage map as lightweight 3D memory evidence. For frame selection, COVRAG maximizes residual coverage gain, iteratively retrieving frames that explain target-view regions not covered by the current context or previously selected memories. To improve scalability in long-video generation, we introduce sliding-window depth caching for efficient geometry estimation. Experiments on RealEstate10K and DL3DV10K show that COVRAG improves long-horizon geometric consistency while maintaining low latency compared to baselines.
Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging. In this work, we revisit the role of positional encoding in video diffusion transformers and show that it provides a useful spatial bias for geometry-aware control. Specifically, if reference tokens are encoded according to their projected locations in the target view, the denoising model is encouraged to retrieve content from position aligned regions of the input video. Building on this observation, we introduce a geometry-aware cross-attention mechanism that enables target video latent tokens to attend to structured context tokens derived from reference images or frames. To establish correspondence between the reference content and the target camera trajectory, we equip the context tokens with a projected positional encoding scheme that combines target-view 2D reprojection with depth-aware disambiguation. At the same time, we preserve the original spatiotemporal positional encoding of the generated video latent, allowing geometric guidance to be injected while maintaining consistency with the video model's native latent structure. The resulting framework provides a simple and effective approach for controllable video generation. It improves spatial controllability in viewpoint-dependent editing tasks, including camera re-trajectory, novel-view video synthesis, and geometry-aware video editing, while preserving the generative prior of the underlying video diffusion model. The code is available at: https://github.com/MTLab/PE-Field.
Yunpeng Bai, Haoxiang Li, Qixing Huang
University of Texas at Austin, USA · Pixocial Technology, USA
Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme scarcity of synchronized, multi-view real-world video data. Consequently, the prevailing paradigm often exhibits limited generalization when processing out-of-distribution real-world videos, with models struggling to accurately adhere to physical scales and camera trajectories. To bridge this gap, we propose Geo-Align, the first Reinforcement Learning framework specifically designed for camera-controlled video re-rendering. Built upon a pretrained model, we optimize the model through a scale-aware perceptual reward mechanism. Specifically, we introduce a metric 3D estimator to extract precise camera trajectories from generated videos, explicitly penalizing deviations in rotation and translation. Furthermore, we meticulously designed a data pipeline strategy based on real-world conditioning videos and target camera trajectories derived from synthetic data, eliminating the reliance on paired data. Extensive experiments demonstrate that Geo-Align consistently outperforms existing supervised learning baselines in both precise camera controllability and visual fidelity, indicating the effectiveness of our method.