FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation
Authors: Yiyi Cai, Yuhan Wu, Kunhang Li, Tu Fangyuan, Xiangyue Zhang, Qiaoge Li, Zhixiang Wang, Kaipeng Zhang, +1 more
Organizations: Alaya Lab · The University of Tokyo · Japan Advanced Institute of Science and Technology · China Mobile Communications Company Limited Research Institute
We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming motion generation model. While FD1 produces plausible motion, it suffers from low efficiency and limited controllability, as its attention design requires repeated computation over the entire history, and it lacks precise trajectory control for real-world applications. To address these limitations and improve generation quality, FD2 introduces three advances. First, Partial Attention makes finalized history representations independent of the active window, enabling KV-cached inference and shared-history packing for efficient training. Second, we establish a necessary-and-sufficient Bregman criterion for regression losses to preserve diffusion's conditional-mean velocity field. This criterion guides an FK-induced quadratic loss that incorporates motion geometry without online FK evaluation. Third, FD2 introduces precise path conditioning to control the character's root trajectory while preserving natural body motion. Experiments show that FD2 reduces training computation by 4.6× and accelerates denoising by 11.29×, reaching 2.303 ms per update on long sequences. Alongside these efficiency gains, FD2 improves motion quality over FD1 and achieves state-of-the-art FID scores among streaming methods, with 0.048 on SEED and 0.053 on HumanML3D.
Figures & tables
Figure 1: FloodDiffusion 2 generates continuous human motion from time-varying text prompts while following a specified path. The character transitions smoothly from walking to running and then crouched walking. Blue and red curves denote the target and generated center-of-mass trajectories, respectively; numbers indicate the generated planar speed in m/s.
Figure 2: Frame-staggered diffusion schedule retained from FD1. (a) Time-shifted denoising ramps separate finalized history, the active window, and future noise. (b) Overlapping active windows advance as frames finish denoising one by one. The illustration uses ns=4 ; β denotes the noise amplitude.
Figure 3: Partial Attention enables KV caching and shared-history training. (a) History queries read their causal prefixes, while active queries read the complete history and interact bidirectionally within the active window. (b) Newly finalized frames are re-encoded causally before their keys and values are appended to the cache. (c) Isolated noisy windows share one causal clean reference, reducing repeated history computation.
Figure 4: Diffusion-compatible losses and the FK-induced quadratic objective. (a) Under the stated regularity and integrability conditions, target-first Bregman divergences with strictly convex potentials, up to a target-only term, are necessary and sufficient for uniquely recovering the conditional mean. (b) Kinematic response Gram matrices estimated from training data are averaged and normalized, then combined with an identity term to obtain a fixed positive-definite metric, avoiding online FK evaluations.
Attention
R@3 ↑
FID ↓
MM-Dist ↓
Div. →
Full
0.810
0.057
2.887
9.579
Causal
0.793
0.096
3.005
9.369
Partial
0.814
0.058
2.839
9.438
Table 1: Attention-mask ablation on HumanML3D with the same 263-D configuration, checkpoint budget, and K=1 . Diversity closer to real motion (9.503) is better.
Full, uncached
Partial, KV cached
Cache
Speedup
History
GFLOPs ↓
ms ↓
GFLOPs ↓
ms ↓
MiB
×
0
6.26
1.810
6.26
1.831
0
0.99
90
23.26
2.107
6.53
1.953
2.8
1.08
450
96.61
3.066
6.90
1.962
14.1
1.56
4500
1506.95
26.002
11.01
2.303
140.6
11.29
Table 3: KV-cache inference on RTX 4090: batch 1, 30 active frames, 16 cached text tokens, CUDA Graph, PyTorch SDPA, no CFG. FLOPs and latency include the complete backbone update, causal re-encoding, and K/V writes for the newly finalized frame. Cache counts bf16 history K/V.
Figure 5: Prompt switching along a commanded path. Body colors indicate the active prompt. Blue is the commanded center-of-mass path, red the generated center-of-mass track; numbers show planar speed in m/s. The examples include repeated jumping and backward walking.
Figure 6: Speed control under fixed text prompts. The walking (a) and running (b) sequences follow varying commanded speed profiles. Body and trajectory colors indicate generated planar speed, illustrating gait adaptation along the commanded path while the action prompt remains unchanged.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: HumanML3D-calibrated FK matrices. Trace-normalized kinematic-response (263-D) and surface-pullback (138-D) matrices with descending eigenvalue spectra. The 263-D heatmap shows the active 67-channel block; omitted channels have zero weight. Each heatmap uses its own signed color scale.
Figure 8: Sampling under three standalone regression losses. (a) Cumulative distributions of the decoded rotation vector projected onto (1,2,3)⊤/14 . (b) Probability assigned to each data region, determined by the nearest training rotation in FK output space. Curves and bars show means over two training seeds; whiskers show their ranges. MSE and the FK-induced quadratic retain all eight regions, whereas nonlinear FK concentrates samples in the highest-density region.
Figure 9: Biomechanical center-of-mass mapping to the neutral SMPL-H body. A front torso detail and a full-body side view locate the anatomical landmarks and the mass-weighted whole-body center of mass. The table compares the landmark heights from de Leva (1996) with their SMPL counterparts. The ground root is the planar projection of the resulting center of mass.
SMPL joint
Weight
SMPL joint
Weight
Pelvis
0.05846
Neck
0.07875
L_Hip
0.11777
L_Collar
0.00000
R_Hip
0.11777
R_Collar
0.00000
Spine1
0.05846
Head
0.06940
L_Knee
0.08225
L_Shoulder
0.01146
R_Knee
0.08225
R_Shoulder
0.01146
Appendix
Table 7: SMPL-H node weights for the ground-projected root pivot. Weights follow the segment masses and center locations in de Leva (1996) , using the stated landmark mapping and normalization to unit sum.
Root definition
Jerk ↓ ( m/s3 )
Curvature ↓ ( m−1 )
Pelvis projection (263-style)
33.450
32.632
Mass-weighted center-of-mass projection (ours)
26.650
29.504
Appendix
Table 8: Smoothness of the global root definition. Statistics are computed over the complete HumanML3D training split at 30 fps, with each motion weighted equally. Jerk is the mean norm of the third temporal derivative of planar position; curvature is evaluated at frames whose planar speed exceeds 0.05m/s .
Representation
Dim.
Pelvis error ( ∘ ) ↓
Body-local error ( ∘ ) ↓
HumanML3D ( Guo et al., 2022a )
263
1.635
30.680
MotionStreamer ( Xiao et al., 2025 )
272
<10−5
<10−5
FloodDiffusion 2 (ours, 138-D)
138
<10−5
<10−5
Appendix
Table 9: Source-rotation round trip. Frame-weighted mean geodesic error over the HumanML3D source-motion collection (26,846 motions; two SVD failures excluded for 263-D).
Dataset
Variant
Predicted channels
Checkpoint step
CFG scale
HumanML3D
FD2
263
60K
4
HumanML3D
FD2-Path
260
55K
3
SEED
FD2
138
300K
2
SEED
FD2-Path
135
300K
2
Appendix
Table 10: Dataset-specific checkpoint and inference settings. The number of predicted channels excludes the three clean root channels for FD2-Path.
We introduce Triangular Resampling (TR), a post-training method for mitigating long-horizon error accumulation in motion diffusion models. Built on FloodDiffusion's triangular denoising schedule, TR addresses the mismatch between ground-truth-derived training windows and model-generated inference states. Replacing only completed motion history leaves this mismatch unresolved in partially denoised states within the active window. TR therefore extends rollout-based training to these states, using ground-truth clamping to limit excessive drift. For each replayed sample, TR draws one denoising threshold, shared across latent positions and replay updates, and replays multi-step triangular denoising without gradient tracking. After each update, states below the threshold are replaced with noise-matched ground truth, while those at or above it retain model predictions. The resulting latent window enters the standard training update. This rollout construction supports both supervised training (TR) and distribution matching (TR-DMD). On 120-second motion generation from HumanML3D test prompts, TR and TR-DMD achieve state-of-the-art FID AUC within their respective non-DMD and DMD comparison groups. Supervised TR reduces FID AUC by 40.9% and FID degradation slope by 55.3% relative to matched post-training without replay.
Kunhang Li, Yiyi Cai, Xiangyue Zhang +5
The University of Tokyo · Alaya Lab · Japan Advanced Institute of Science and Technology
Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous video stream. However, streaming video diffusion models introduce a fundamental train-inference mismatch: inference follows a specialized denoising order, whereas advanced training strategies typically require diverse noise-level configurations. To address this trade-off between train-inference consistency and training coverage, we reformulate the video diffusion sampling as a frame-indexed stochastic process over noise levels. Within this stochastic process space, we construct a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling. We further introduce a joint calibration algorithm and a temporal correlative sampling algorithm to ensure trajectory smoothness and cross-frame correlation. Building on these designs, we propose Stream Forcing, a unified training framework for streaming video generation that balances training sufficiency and inference efficiency. Extensive experiments demonstrate that Stream Forcing significantly improves generation quality with a 36.6% FVD improvement on the UCF-101 benchmark. Furthermore, our method facilitates robust zero-shot extrapolation to long-horizon video generation with a 27.9% FVD improvement on the UCF-101 benchmark.
Yueting Zhu, Yuehao Song, Kaicheng Zhang +5
Huazhong University of Science & Technology · Anyverse Dynamics · Horizon Robotics
Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, current motion diffusion models face two major limitations: a representational gap caused by pre-trained text encoders that lack motion-specific information, and error propagation during the iterative denoising process. This paper introduces Reconstruction-Anchored Diffusion Model (RAM) to address these challenges. First, RAM leverages a motion latent space as intermediate supervision for text-to-motion generation. To this end, RAM co-trains a motion reconstruction branch with two key objective functions: self-regularization to enhance the discrimination of the motion space and motion-centric latent alignment to enable accurate mapping from text to the motion latent space. Second, we propose Reconstructive Error Guidance (REG), a testing-stage guidance mechanism that exploits the motion diffusion model's inherent self-correction ability to mitigate error propagation. At each denoising step, REG uses the motion reconstruction branch to reconstruct the previous estimate, reproducing the prior error patterns. By amplifying the residual between the current prediction and the reconstructed estimate, REG highlights the improvements in the current prediction. Extensive experiments demonstrate that RAM achieves significant improvements and state-of-the-art performance. Our code will be released.
Yifei Liu, Changxing Ding, Ling Guo +2
South China University of Technology · Joy Future Academy