cs.CVJun 20, 2026

Feed-forward Motion In-betweening for Any 4D

Authors: Hiroki NishizawaHubert P. H. ShumYoshihiro FukuharaHirokatsu KataokaShigeo Morishima

Organizations: Waseda University · AIST · Durham University

Abstract

4D dynamics (3D geometry evolving over time) is a fundamental representation of the physical world and plays a crucial role in world modeling (e.g., animation and games). Owing to the scarcity of large-scale, long-horizon 4D mesh data with arbitrary shapes, early text-to-4D methods rely on distillation or test-time optimization from video diffusion priors, making inference prohibitively slow. Recent feed-forward generators greatly reduce inference cost but offer limited spatiotemporal controllability, and short-horizon generation often leads to error accumulation in long-horizon sequences. We propose a novel feed-forward in-betweening framework for arbitrary 4D meshes with keyframe conditioning. Building on universal mesh-animation latents, we introduce a frame-wise mesh VAE that encodes each frame into topology-agnostic latent tokens anchored by a reference mesh for keyframe conditioning. We further introduce a keyframe-conditioned rectified flow model with an MMDiT backbone that synthesizes non-keyframe frames conditioned on sparse keyframes. Experiments show strong performance and improved controllability on both DyMesh16 and DyMesh32 benchmarks.

Explore similar work

CardsList
  1. Fast 4D Mesh Generation by Spatio-Temporal Attention Chains

    May 19, 2026Dvir Samuel, Yuval Atzmon, Gal Chechik +1Meshes

  2. Helix4D: Complex 4D Mesh Generation

    May 25, 2026Jiraphon Yenphraphai, Jianqi Chen, Jian Wang +64D GenerationMeshes