cs.CVOct 1, 2026

Form and Void: Entangled Composition through an Autonomous AI Agent

Authors: Shiwen Wang, Jian Yang, Xu Wang, Xincan Wang, Weiming Dong

Organizations: School of AI, University of Chinese Academy of Sciences · Renmin University of China · Shanghai Theatre Academy · MAIS, Institute of Automation, Chinese Academy of Sciences

Abstract

Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbf{F}orm \textbf{a}nd \textbf{V}oid \textbf{A}gent (\textbf{FaV-A}), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.

Figures & tables

Explore similar work

CardsList
  1. Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm

    May 12, 2026Yaofang Liu, Kangning Cui, Meng Chu +7Visual GenerationText Analysis and Detection

  2. IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

    Jun 23, 2026Zixuan Li, Haokun Lin, Yicheng Xiao +10Multi-Reference Image GenerationText-To-Image

  3. ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

    Aug 5, 2026Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5Image GenerationMultimodal Generation