Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows with the generated history. Existing compression strategies discard history using fixed windows or select tokens through local attention and similarity signals, without directly measuring whether a chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find that tokens with larger discrepancies between intermediate clean predictions and final denoised values tend to carry visual evidence less predictable from the retained context. DeCoPrune uses this model-intrinsic signal to retain high-discrepancy tokens in the long-term cache while pruning low-discrepancy tokens. To evaluate information retention, we introduce CMBench, comprising 58 approximately one-minute generated or real-world context episodes and 116 Reappear or Revisit continuation tasks requiring recall of earlier events or objects. Experiments with LingBot World v2 show that DeCoPrune achieves a DINO score of 0.6701 on a 0-1 scale, with an 85.43% reduction in cumulative historical KV token counts and a 4.14-fold continuation-generation speedup over FullKV. Its head-specialized variant reaches 0.6783 at an 86.19% pruning ratio, approaching FullKV's 0.6803 score and exceeding the evaluated compression baselines at similar budgets. These results indicate that denoising consistency can support long-range information retention while reducing autoregressive inference cost. Our project homepage is https://decoprune.github.io. The code is available at https://github.com/DeCoPrune/CMBench, and the benchmark at https://huggingface.co/datasets/Aoraku/CMBench.
Figures & tables
Figure 1: Overview of CMBench and DeCoPrune . (a) A representative CMBench episode showing which historical tokens DeCoPrune retains and reuses during continuation. (b) Consistency–pruning-ratio trade-off across pruning methods. At comparable pruning ratios, our method achieves higher DINO consistency than the evaluated compression baselines.
Figure 2: Denoising consistency and context redundancy. (a) At an intermediate denoising state t=t′ , a local step-to-final extrapolation (red dashed) can deviate from the final denoised state without context ( A→B ). Relevant context concentrates the conditional outcome, aligning the extrapolation more closely with the final state ( C→D ). Our policy treats content that is consistently predictable from existing context as redundant, pruning low-discrepancy tokens while retaining high-discrepancy tokens. (b) Curves report the MSE between the x0 prediction at each intermediate denoising step and the final x0 . The Causal Forcing, LongLive, and Self-Forcing plots compare runs with and without the available context KV: first-frame KV for Causal Forcing and Self-Forcing, and preceding-chunk KV for LongLive. For LingBot World v2, which does not support text-to-video generation, both runs use an initial frame and rotate the camera: “with context” means that the content revealed after rotation appeared earlier in the context, whereas “without context” means that it did not. In all four comparisons, the with-context condition has lower error across the plotted steps, consistent with our criterion.
Figure 3: DeCoPrune overview. A probe and final prediction from the same denoising trajectory yield a token-retention mask. The mask physically compacts historical KVs after the recent-window delay.
Figure 4: Overview of CMBench . Top: a one-minute generated context assembled from six prompted clips; three target events yield three continuation tasks. Bottom: the corresponding continuation prompts and evaluation pipeline. The reference target and its generated counterpart are localized and segmented, then compared using DINO similarity.
Method
DINO ↑
PR ↑
FPS ↑
Speedup ↑
Temporal
Motion
Aesthetic
Image
Flickering ↑
Smoothness ↑
Quality ↑
Quality ↑
FullKV
0.6803
0.00%
1.568
1.00 ×
0.9499
0.9705
0.4633
0.7125
Streaming ( Xu et al., 2026c ; Yang et al., 2025 )
0.4592
94.37%
9.591
6.12 ×
0.9474
0.9698
0.4769
0.7103
DummyForcing ( Guo et al., 2026 )
0.4461
98.03%
7.393
4.72 ×
0.9425
0.9687
0.4677
0.6933
ForcingKV ( Ji et al., 2026 )
0.5313
80.62%
5.288
3.37 ×
0.9437
0.9661
0.4610
0.7018
TempDiff ( Hwang et al., 2024 ; Fu et al., 2025 )
0.6229
86.36%
6.443
4.11 ×
0.9425
0.9634
0.4677
0.7161
Table 1: Main comparison on CMBench and VBench . DINO is reported on a 0–1 scale; results are averaged over three random seeds. FPS and speedup over FullKV measure continuation generation only, excluding prefix processing.
Figure 5: Qualitative comparison under the matched settings used in Tables 1 and 3 . We highly recommend viewing the video comparisons on the supplementary webpage to better appreciate temporal consistency and visual detail.
Figure 6: Ablation and metric analysis on 13 independent cases. (a) Threshold γ is swept from light to dark at each probe-step index; the marked operating point is s∗=2 ( τ∗=899 ), γ=0.10 . (b) VBench subject/background consistency and CMBench DINO versus PR.
Table 8
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Effect of RoPE re-indexing across context positions. The heatmaps show the distribution of CMBench cases over context-prompt time and DINO similarity for DeCoPrune–HS (top) and FullKV (bottom), without re-indexing (left) and with re-indexing (right). Cell values and color intensity indicate the number of cases. In the separate diagnostic evaluation of Table 3 , re-indexing shifts the score distribution upward for both methods, increasing the mean DINO similarity from 0.6686 to 0.7156 for DeCoPrune–HS and from 0.4385 to 0.6804 for FullKV .
Figure 8: Full-context comparison across autoregressive video backbones. We compare ground-truth reference frames with FullKV generations from LingBot World v2 , Causal Forcing , Self-Forcing , and LongLive on three representative object-retrieval cases that require no camera-control input. Despite retaining the complete context, the three alternative backbones frequently fail to recover the target object, whereas the LingBot World v2 generations more closely reproduce the targets in these examples.
Backbone
Mean DINO ↑
LingBot World v2 ( Gao et al., 2026 )
0.6803
Causal Forcing ( Zhu et al., 2026 )
0.2627
Self-Forcing ( Huang et al., 2025a )
0.2609
LongLive ( Yang et al., 2025 )
0.2608
Appendix
Table 4: Full-context performance of candidate backbones on CMBench . Mean DINO (0–1) is computed over the same completed subset for all backbones.
Method
PR ↑
DINO ↑
FullKV
0.00%
0.6106
DeCoPrune
57.25%
0.5880
Streaming
53.57%
0.5162
ForcingKV
54.17%
0.5302
DummyForcing
55.60%
0.5157
Appendix
Table 5: LongLive continuation with 10-second contexts on 18 cases. DINO is reported on a 0–1 scale.
Method
Synthetic Video
Real Video
FullKV
0.6873
0.6621
DeCoPrune
0.6909
0.6412
TempDiff ( Hwang et al., 2024 ; Fu et al., 2025 )
0.6484
0.5212
ForcingKV ( Ji et al., 2026 )
0.5422
0.3839
Streaming ( Xu et al., 2026c ; Yang et al., 2025 )
0.4592
0.3274
DummyForcing ( Guo et al., 2026 )
0.4549
0.3111
Appendix
Table 6: Evaluation by context-video source. Mean DINO scores (0–1) on synthetic-video and real-video contexts using one random seed.