Form and Void: Entangled Composition through an Autonomous AI Agent
Organizations: School of AI, University of Chinese Academy of Sciences · Renmin University of China · Shanghai Theatre Academy · MAIS, Institute of Automation, Chinese Academy of Sciences
Abstract
Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbf{F}orm \textbf{a}nd \textbf{V}oid \textbf{A}gent (\textbf{FaV-A}), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.
Figures & tables
| Model Variant | TA | GQ | AS |
|---|---|---|---|
| w/o Topic Analysis | |||
| w/o Detailed Base Prompt | |||
| w/o Image Feature Analysis | |||
| w/o Specific Composition | |||
| w/o Multi-modal Generation | |||
| FaV-A (Ours - Full) | 4.65 0.35 | 4.58 0.40 | 4.52 0.38 |