cs.CVDec 4, 2024

Auto-Regressive Models Need Structural Registers: Semantic Foresight Improves Coherent Driving Video Continuation

Authors: Ruibo Ming, Jingwei Wu, Zhewei Huang, Zhuoxuan Ju, Jianming Hu, Lihui Peng, Shuchang Zhou

Abstract

Front-view driving video continuation is a critical component for constructing sophisticated world models. However, maintaining long-term coherence faces the fundamental challenge of mitigating ''generative degeneration'' during the auto-regressive process. This phenomenon arises when the limited attention capacity of world models becomes saturated with high-entropy pixel details over extended sequences, leading to a collapse of global scene structure and motion intensity. To address this, we introduce STRIDE (STructural RegIsters for Decoupled Extrapolation), a large vision model architecture with an enforced semantic foresight mechanism. STRIDE decouples video continuation into two sub-tasks via interleaved generation of semantic and RGB tokens: forecasting low-entropy semantic maps, and subsequently generating high-frequency visual appearance conditioned on this structure. The semantic tokens explicitly function as ''structural registers'' -- dedicated memory slots that robustly retain the scene's long-term context and world dynamics. Furthermore, we employ an architecturally decoupled flow matching-based super-resolution module to enhance the final visual fidelity. Extensive experiments in autonomous driving scenarios demonstrate that STRIDE effectively mitigates generative degeneration, achieving minute-level, temporally consistent driving video continuation with strong generalization ability.

Explore similar work

CardsList
  1. DriveVA: Video Action Models are Zero-Shot Drivers

    Apr 5, 2026Mengmeng Liu, Diankun Zhang, Jiuming Liu +7Autonomous DrivingDrives

  2. Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency

    Jul 13, 2026Zihan Su, Teng Hu, Jiangning Zhang +4Video World Models