cs.CVMar 31, 2026

TrajectoryMover: Generative Movement of Object Trajectories in Videos

Authors: Kiran ChhatreHyeonho JeongYulia GryaditskayaChristopher E. PetersChun-Hao Paul HuangPaul Guerrero

Abstract

Generative video editing has enabled creative control over an object's position in a video by prescribing an object's 3D or 2D motion trajectory, while preserving both video plausibility and identity. However, manually specifying a plausible full motion trajectory, like the arcs of a bouncing ball, requires time and expertise, and may therefore not be a suitable editing task for non-experts or quick edits. In contrast, in the image domain, generative object translation has been established as a simple editing task that requires only a single drag to move an object to a new location; yet an equivalent method is still missing for videos. We propose an analogous new editing task for videos that translates an object's 3D motion trajectory in a video while preserving identity and plausibility of appearance and motion, for example, translating the trajectory of a bouncing ball while preserving its relative motion. The main challenge in training this task lies in obtaining paired video data for this scenario. We introduce TrajectorySynth, a new data generation strategy for large-scale synthetic paired video data and a video generator TrajectoryMover fine-tuned with this data. We show that this enables generative movement of object trajectories. Project Page: https://chhatrekiran.github.io/trajectorymover

Explore similar work

Aug 1, 2026cs.CV

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or recover motion from fully generated videos through costly, error-prone perception pipelines. We close the supervision gap with the MOVE dataset, 5,038 object-centric egocentric trajectories, each paired with a fine-grained natural-language instruction rather than a coarse verb-noun label. We further propose DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference: rather than generating a video, it reads motion from the internal representations of a frozen image-to-video diffusion model at an early denoising step. A lightweight flow-matching Reader decodes query-key attention tracks and pooled hidden states into relative 6-DoF poses. To our knowledge, this is the first approach to directly decode object 6-DoF trajectories from intermediate video diffusion representations rather than generated pixels. DreamTraj sets a new state of the art on both translation and rotation against forecasters that consume multi-frame or privileged inputs, and runs 4.6x faster than generate-then-extract pipelines.
Tongsheng Ding, Zhen Luo, Yixuan Yang +4
Dec 3, 2025cs.CV

LAMP: Language-Assisted Motion Planning for Controllable Video Generation

Video generation has achieved remarkable progress in visual fidelity and controllability, enabling conditioning on text, layout, or motion. Among these, motion control - specifying object dynamics and camera trajectories - is essential for composing complex, cinematic scenes, yet existing interfaces remain limited. We introduce LAMP that leverages large language models (LLMs) as motion planners to translate natural language descriptions into explicit 3D trajectories for dynamic objects and (relatively defined) cameras. LAMP defines a motion domain-specific language (DSL), inspired by cinematography conventions. By harnessing program synthesis capabilities of LLMs, LAMP generates structured motion programs from natural language, which are deterministically mapped to 3D trajectories. We construct a large-scale procedural dataset pairing natural text descriptions with corresponding motion programs and 3D trajectories. Experiments demonstrate LAMP's improved performance in motion controllability and alignment with user intent compared to state-of-the-art alternatives establishing the first framework for generating both object and camera motions directly from natural language specifications. Code, models and data are available on our project page.
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra +3
Oct 28, 2025cs.CV

VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos

Although recent video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions, which is crucial in creating truly original and artistic videos. The challenge lies in finding sufficient training videos with the intended uncommon camera motions. To this end, we propose VividCam, a training paradigm that enables diffusion models to learn complex camera motions from synthetic videos, releasing the reliance on collecting realistic training videos. VividCam incorporates multiple disentanglement strategies that isolate camera motion learning from synthetic appearance artifacts, ensuring more robust motion representation and mitigating domain shift. We show that our design synthesizes a wide range of precisely controlled camera motions using surprisingly simple synthetic data. Notably, this synthetic data often consists of basic geometries within a low-poly 3D scene and can be efficiently rendered by engines like Unity. Our video results can be found in https://wuqiuche.github.io/VividCamDemoPage/ .
Qiucheng Wu, Handong Zhao, Zhixin Shu +3