Dynamics-Inspired Diffusion for Foreground-Preserving Document Background Editing
Organizations: University of Maryland at College Park, United States
Abstract
We revisit diffusion-based generation for structured visual content and identify a fundamental limitation of existing approaches: foreground preservation and background stylization are typically enforced through external interventions, such as hard masking or corrective post-processing, rather than arising from the generative process itself. Here, we define background as the generative content outside designated foreground regions (e.g., text and layout elements), while preserving the structural integrity of the foreground. We propose a dynamical systems perspective on diffusion, in which controllable generation is formulated as trajectory shaping in latent space. Under this view, we introduce Auxiliary Context Diffusion (ACD), a state-space control framework that integrates heterogeneous signals (layout-derived foreground indicators, document summaries, and style representations) directly into the diffusion dynamics. This formulation induces time-scale separation in the generative process, where foreground regions become dynamically stabilized while background regions remain expressive. To address stylistic drift across multi-page documents, we further introduce style directions as persistent latent constraints that guide diffusion trajectories within a shared stylistic subspace. Unlike prior approaches that entangle style with prompt conditioning, our formulation enables reusable and consistent style control across pages. We validate the proposed perspective through controlled experiments on synthetic document benchmarks, demonstrating that trajectory-level control provides a unified and extensible mechanism for structured generation without retraining, hard masking, or corrective post-processing. These results suggest a new direction for controllable diffusion in document-centric and multimodal applications.
Figures & tables
| Method | Layout | Color | Graphic Style | Compliance | WCAG (%) | OCR Acc. | CLIP MP Consistency | CLIP Prompt Score | LLM Voting |
|---|---|---|---|---|---|---|---|---|---|
| BAGEL | 3.7335 | 3.9735 | 3.8857 | 3.7292 | 82.95 | 0.363 | 0.6317 | 0.1567 | 3.8478 |
| GPT-5 | 4.155 | 4.3272 | 4.321 | 4.2614 | 80.75 | 0.7217 | 0.3113 | 0.1657 | 4.2363 |
| GPT-5 (Naive Overlay) | 3.8928 | 4.0486 | 4.0778 | 4.13 | 56.87 | 0.3077 | 0.6085 | 0.2631 | 3.9943 |
| Ours | 4.24 | 4.07 | 4.14 | 4.74 | 98.12 | 0.779 | 0.6785 | 0.3144 | 4.2992 |
| Ours w/o Style Bank | 4.1878 | 3.9492 | 4.055 | 4.70 | 98.08 | 0.7769 | 0.6661 | 0.3083 | 4.2335 |
| Ours w/o SSC | 3.6885 | 3.7142 | 3.8442 | 4.2785 | 54.20 | 0.333 | 0.6667 | 0.2781 | 3.8764 |
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Example user prompt |
|---|---|
| Geometric | “Add a background with a modern abstract design composed of layered geometric forms, featuring clean repetitions and harmonious symmetry for a structured visual effect.” |
| Shapes | “Add a background featuring a playful yet balanced composition of varied shapes, combining bold curves and soft angles to create natural depth.” |
| Textures | “Add a background with richly layered textures, where smooth and coarse surfaces interact to produce tactile depth and visual interest.” |
| Colorful | “Add a background with a vivid, lifelike scene filled with diverse colors under natural lighting, creating a vibrant and dynamic atmosphere.” |
| Muted | “Add a background with a softly lit, realistic setting using a desaturated color palette that conveys calmness and understated elegance.” |
| Professional | “Add a background with a refined and realistic design, emphasizing clean lines, minimal clutter, and subtle details for a polished appearance.” |
| Method category (representatives) | Operates on | Foreground | Multi-page | Dense text | Training- |
|---|---|---|---|---|---|
| existing pages | control mechanism | consistency | preservation | free | |
| Layout / poster generation [ 5 , 47 , 32 , 44 , 40 ] | ✗ (regenerates layout) | via re-layout | ✗ | ✗ | — |
| Mask-guided inpainting [ 19 , 26 , 1 , 9 ] | partial | binary spatial mask | ✗ | ✗ | ✓ |
| Text-aware diffusion [ 6 , 7 , 29 ] | partial | assumes blank / rendered text | ✗ | ✗ | — |
| Layered diffusion [ 17 , 23 , 45 ] | ✗ (requires clean layers) | layer separation | ✗ | ✗ | ✗ |
| Attention control [ 14 , 3 , 4 ] | partial | attention re-weighting | ✗ | ✗ | ✓ |
| Axis | Kang et al. [ 20 ] | Ours (ACD) |
|---|---|---|
| Foreground preservation | fixed attenuation mask | time-dependent trajectory gating (Eq. 9 ) |
| Readability | post-hoc semi-transparent backing shapes | stabilized canvas alone (WCAG 98.12%) |
| Cross-page consistency | per-page LLM-generated instructions | persistent latent style direction |
| Per-page inference dependency | one LLM call per page | none |
| Method (3-page document) | Peak GPU memory (GB) | Wall-clock time (sec) |
|---|---|---|
| BAGEL | 31.3461 GB / 40.000 GB | 246.00 sec |
| GPT-5 (hosted API) | N/A | 324.00 sec |
| Ours w/o SSC | 29.1953 GB / 40.000 GB | 243.00 sec |
| Ours w/o Style Bank | 29.1953 GB / 40.000 GB | 243.00 sec |
| Ours (full ACD) | 29.1953 GB / 40.000 GB | 243.00 sec |