cs.CVOct 1, 2026

SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing

Authors: Fa-Ting Hong, Peter Wonka

Organizations: King Abdullah University of Science and Technology

Abstract

Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbf{Source-Preserving Denoising framework (SuperMotion)}, which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Omni-Supervised Motion Editing: Balancing Change and Invariance through Positive-Negative Learning

    May 29, 2026Zhenwu Shi, Jingyu Gong, Peiwei Wang +7Text-To-Motion GenerationImage-Text Alignment

  2. UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

    Aug 10, 2026Yilei Hua, Beibei Jing, Ce Zheng +3Human Motion GenerationText-To-Motion Generation