cs.CVMay 12, 2026

hh-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement

Authors: Yuzhu WangXi YeDuo SuYangyang XuJun Zhu

Organizations: Department of Computer Science and Technology, Tsinghua University · South China University of Technology

Abstract

Training-free camera control for pretrained flow-matching video generators is a partial-observation inverse problem: a depth-warped guidance video supplies noisy evidence on a subset of latent sites, which the sampler must reconcile with the pretrained prior. Existing methods struggle to balance the trade-off between trajectory adherence and visual quality and the heuristic guidance-strength tuning lacks robustness. We propose \textbf{hh-control}, which resolves this dilemma through a structural change to the sampler: each outer hard-replacement guidance step is augmented with an inner-loop \emph{block-conditional pseudo-Gibbs refinement} on the unobserved complement at the same noise level, with provable convergence to the partial-observation conditional data law. To accelerate convergence on high-dimensional video latents, we exploit their conditional locality, partitioning the unobserved complement into 3D patches, each tracked by a custom mixing indicator that adaptively freezes converged patches. On RealEstate10K and DAVIS, \textbf{hh-control} attains the best FVD against all seven training-free and training-based competitors, outperforming every training-free baseline on every reported metric.

Explore similar work

CardsList
  1. GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

    Dec 9, 2025Frédéric Fortier-Chouinard, Yannick Hold-Geoffroy, Valentin Deschaintre +2Controllable Video GenerationCamera Control