cs.CVJan 29, 2026

Dynamics-Inspired Diffusion for Foreground-Preserving Document Background Editing

Authors: Taewon Kang, Yu Shen, Ming C. Lin

Organizations: University of Maryland at College Park, United States

Abstract

We revisit diffusion-based generation for structured visual content and identify a fundamental limitation of existing approaches: foreground preservation and background stylization are typically enforced through external interventions, such as hard masking or corrective post-processing, rather than arising from the generative process itself. Here, we define background as the generative content outside designated foreground regions (e.g., text and layout elements), while preserving the structural integrity of the foreground. We propose a dynamical systems perspective on diffusion, in which controllable generation is formulated as trajectory shaping in latent space. Under this view, we introduce Auxiliary Context Diffusion (ACD), a state-space control framework that integrates heterogeneous signals (layout-derived foreground indicators, document summaries, and style representations) directly into the diffusion dynamics. This formulation induces time-scale separation in the generative process, where foreground regions become dynamically stabilized while background regions remain expressive. To address stylistic drift across multi-page documents, we further introduce style directions as persistent latent constraints that guide diffusion trajectories within a shared stylistic subspace. Unlike prior approaches that entangle style with prompt conditioning, our formulation enables reusable and consistent style control across pages. We validate the proposed perspective through controlled experiments on synthetic document benchmarks, demonstrating that trajectory-level control provides a unified and extensible mechanism for structured generation without retraining, hard masking, or corrective post-processing. These results suggest a new direction for controllable diffusion in document-centric and multimodal applications.

Figures & tables

Appendix figures & tables37 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Text-Conditioned Background Generation for Editable Multi-Layer Documents

    Dec 19, 2025Taewon Kang, Joseph K J, Chris Tensmeyer +4Precise EditingHigh-Fidelity Conditional Generation

  2. UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement

    Apr 20, 2026Jingwei Yang, Ruoxi Wu, Wei Shen +4Photorealistic Style TransferStyle