cs.CVSep 30, 2026

Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video

Authors: Chang-Bae Bang, Hyungjin Chung, Byung-Hoon Kim

Organizations: Department of Psychiatry, Yonsei University College of Medicine · Institute of Behavioral Sciences in Medicine, Yonsei University College of Medicine · Department of Biomedical Systems Informatics, Yonsei University College of Medicine · Department of Computer Science & Engineering, Korea University · Yonsei Institute for Digital Healthcare, Yonsei University

Abstract

Studying the alignment between the internal representations of vision models and the responses of the visual cortex to the same observed visual stimuli has enabled us to better understand human visual processing. However, studies so far have largely overlooked the fact that the human brain not only processes observed visual stimuli, but also predicts upcoming stimuli based on what has been observed. Accordingly, we hypothesize that internal representations for generating future video frames are better aligned with the predictive nature of human visual processing than representations of the observed video itself. To this end, we compare the alignment between human video-watching fMRI responses in the visual cortex and the internal representations from two types of video diffusion models, an autoregressive (AR) model and its non-AR base model. We first conduct a within-model analysis of the AR video diffusion model and show that the representations for future video generation align better with the visual cortex than the representations of the observed video. We then compare the internal representations of the AR model with those of its non-AR base model and again show that the representations for future video generation align better with the visual cortex than the representations for observed video reconstruction by the base model. Specifically, the alignment of observed video reconstruction is concentrated in lower-order visual cortex, whereas that of future video generation is concentrated in higher-order visual cortex. Finally, we show in a human behavioral experiment that humans prefer videos generated by amplifying the contributions of individual layers that align better with the visual cortex.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Video-Mirai: Autoregressive Video Diffusion Models Need Foresight

    Jun 2, 2026Yonghao Yu, Lang Huang, Runyi Li +2Autoregressive Video Diffusion ModelsDiffusion Models

  2. Video Generation Models: A Survey of Post-Training and Alignment

    Sep 30, 2026Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni +10Generative Video ModelsVideo Generation

  3. YoCausal: How Far is Video Generation from World Model? A Causality Perspective

    May 28, 2026You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee +3Video Diffusion ModelsCausal Reasoning