Research
Form and Void: Entangled Composition through an Autonomous AI Agent
Overview Research area: Computer vision and multimodal generative AI, specifically MLLM-based agents for design-oriented image generation (positive-negative space composition). Technical level: Interm

- arXiv
- 2610.02045
- Published
- 2026-10-01
- Authors
- Shiwen Wang, Jian Yang, Xu Wang, Xincan Wang, Weiming Dong
AI summary
Overview
Research area: Computer vision and multimodal generative AI, specifically MLLM-based agents for design-oriented image generation (positive-negative space composition).
Technical level: Intermediate. The method is a prompt-driven, staged agent pipeline built on existing multimodal models rather than a newly trained architecture, but it assumes familiarity with MLLMs, diffusion-based text-to-image generation, in-context learning, and figure-ground design concepts.
Scope: The paper presents Form and Void Agent (FaV-A), a three-stage autonomous multimodal agent that turns an abstract text topic into a positive-negative space artwork in which a foreground form and a negative-space form share a single continuous contour.
What This Paper Is About
Positive-negative space composition asks a designer to make the empty region around a subject carry a second, equally meaningful shape, with both shapes sharing one boundary. Directly prompting today's text-to-image models or multimodal large language models (MLLMs) to do this in a single pass tends to produce structurally unstable or semantically ambiguous images. The paper's goal is to show that splitting the task into successive stages — generate a base object, analyze its geometry, then synthesize the entangled final image under those constraints — produces more coherent results than zero-shot single-pass prompting.
Key Contributions
- Task formulation. The authors frame positive-negative space generation as a multimodal design problem that requires coordinated spatial and semantic reasoning under a shared-boundary constraint, rather than as ordinary text-to-image synthesis.
- The FaV-A agent. They propose a staged multimodal agent that decomposes the task into three successive phases: base-object generation, image-aware structural analysis, and instruction-guided final synthesis. The pipeline is explicitly mask-free, in contrast to layout-controlled generation methods that rely on manually specified masks, edges, or spatial layouts.
- Empirical validation. Through qualitative comparison against three closed-source baselines (Gemini, GPT, Claude), a human user study with N = 50 participants rating five ablated variants on a 5-point Likert scale, and a visual ablation study, the authors report that the staged framework produces more structurally coherent and semantically aligned compositions than direct zero-shot MLLM prompting.
- A taxonomy of baseline failure modes. The paper characterizes three distinct failure patterns in direct-prompting models: semantic drift, additive composition, and graphic plausibility without true figure-ground reversal.
Main Findings
- Direct prompting fails to entangle shapes. Given the same topic and explicit positive-negative space design knowledge, the three baselines each failed in a different way (Figure 3). Claude showed the weakest alignment with the intended topics, with "few clearly recognizable topic-related elements" remaining in cases such as Rose and Love and War and Peace. Gemini preserved topics but used additive composition — stacking, juxtaposing, or locally inserting one object into another — so its forms stayed "illustrative rather than structurally interdependent." GPT showed the strongest design sensibility of the three direct-prompting baselines, producing clearer visual hierarchy and better poster-like readability, but the authors characterize its outputs as "literal graphic design" rather than genuine figure-ground reversal.
- FaV-A achieves shared-contour entanglement. The authors attribute this to converting the illusion into a progressive, constrained reasoning task: a base image establishes geometry, multi-dimensional parsing produces a semantic spatial anchor (P_anchor), and that anchor maps semantics onto specific topological regions, forcing positive and negative space to share one continuous contour.
- Ablation with human evaluation. Table 1 reports means and standard deviations over N = 50 raters on Topic Adherence (TA), Gestalt Quality (GQ), and overall Aesthetic Score (AS):
- w/o Topic Analysis: TA 1.35 ± 0.42, GQ 3.12 ± 0.65, AS 2.85 ± 0.70
- w/o Detailed Base Prompt: TA 4.15 ± 0.50, GQ 2.90 ± 0.58, AS 3.25 ± 0.62
- w/o Image Feature Analysis: TA 3.80 ± 0.55, GQ 2.15 ± 0.60, AS 2.40 ± 0.65
- w/o Specific Composition: TA 3.92 ± 0.48, GQ 1.85 ± 0.52, AS 2.10 ± 0.55
- w/o Multi-modal Generation: TA 3.45 ± 0.60, GQ 1.60 ± 0.45, AS 1.80 ± 0.48
- FaV-A (full): TA 4.65 ± 0.35, GQ 4.58 ± 0.40, AS 4.52 ± 0.38
- Removing topic analysis destroys semantics. The largest drop in Topic Adherence came from bypassing semantic planning (TA 1.35), which the authors read as the model's inability to follow the semantic prompt. In the ablation figures, this degeneration appears as shallow vertical juxtaposition of elements.
- Compositional constraints and multimodal generation govern Gestalt quality. The heaviest Gestalt Quality penalties came from removing specific compositional constraints (GQ 1.85) and bypassing multi-modal generation (GQ 1.60). Without the final multimodal reasoning pass, the visual illusion collapses entirely — in one example the cat is arbitrarily transformed into a dog.
- Underspecified base prompts cause early semantic drift. Replacing the LLM-derived base object prompt with a generic manual prompt ("Positive or negative shapes representing {topic}") generated an irrelevant dragonfly-like shape instead of a player, and this error propagated through the rest of the pipeline.
- Removing image feature analysis breaks structural fidelity. Without multi-dimensional parsing of the base image, the refinement process disconnects from the original object: the player's hair characteristics are noticeably altered and the cat's pose drifts. The authors note this matters especially for human-in-the-loop workflows where a designer's chosen base form must stay morphologically stable.
- Zero-shot topological and cross-domain behavior. Qualitative examples include Silent Duel: Between the Pieces and Their Masters, where negative space around a chess piece is sculpted into a human profile; Phantom of Peace: A Doomsday Mushroom Cloud Beneath the Olive Branch, where the mushroom cloud is formed through the absence of surrounding textures rather than explicit contours; Whispered Kiss Beneath the Rose; Harmonious Gaze: Poetry of Cat and Cat/Dog in Silent Companionship; and a CAT/DOG typography series that fuses text characters with animal silhouettes into bistable configurations.
- No quantitative baseline comparison is reported. The paper reports numbers only for the ablation variants and the full pipeline in Table 1. The comparison against Gemini, GPT, and Claude is presented qualitatively in Figure 3; a quantitative head-to-head user study against those baselines is not reported.
- No dataset is reported. Evaluation uses abstract topics ("Chess and Player," "Petting a Cat," and the topics named in the baseline discussion) rather than a benchmark dataset; the paper does not report a dataset size or a standard benchmark.
Methodology in Plain English
The system does not try to generate the finished illusion in one shot. Instead it runs a three-stage conversation-like workflow on top of a large multimodal model, denoted as M, plus a text-to-image module.
Stage I — Semantic planning. Given an abstract topic T, a planning model is given a system prompt containing design priors that emphasize figure-ground structure, negative space, and compositional balance. It outputs two things: a Base Object Prompt describing the primary subject, and a Base Composition Blueprint describing the intended figure-ground arrangement. This turns a vague topic into something concrete enough to draw.
Stage II — Spatial anchoring. The base object prompt is sent to a text-to-image module, producing a base image. The agent then examines that image along several dimensions: shape attributes, spatial relationships, visual hierarchy, potential negative-space interpretations, symbolic associations, and implementation-related considerations. To align this analysis with the target, the agent also receives three visual-textual exemplars (N = 3) plus the composition blueprint. The output is a detailed compositional prompt — the "anchor" — that describes intended contour relationships, the negative-space interpretation, and compositional constraints in natural language rather than as explicit spatial masks.
Stage III — Guided final synthesis. A synthesis model takes the base image plus the anchor prompt and produces the final image together with a textual description of the semantic relationship it expresses. Because the anchor specifies which contours should align with what, the final stage can force the two interpretations onto a single continuous edge.
The key design decision is that no masks, bounding boxes, or edge-control signals are supplied by hand. Structure is discovered from the model's own intermediate image and then re-expressed as text guidance.
Why This Matters
Impact on research. The paper argues that multimodal agents can serve as "autonomous cognitive partners in high-level aesthetic design," and it positions staged intermediate reasoning as an alternative to both explicit spatial-control pipelines (ControlNet, T2I-Adapter, GLIGEN) and prior design agents (PosterLLaVA, GenArtist, BannerAgency, GraphicBench). It also frames the results as evidence of zero-shot topological reasoning in MLLMs — a capability distinct from the layout planning and tool orchestration that prior design agents target.
Real-world applications (as suggested by the work's framing):
- Logo and brand mark design, where a single contour must carry two meanings at once.
- Poster, editorial, and print graphic design that relies on figure-ground ambiguity and high-contrast black-and-white composition.
- Typography and lettering, demonstrated by the CAT/DOG typography series that fuses text characters with animal silhouettes.
- Human-in-the-loop creative tooling, where a designer picks an intermediate base form and the system preserves its morphology during later refinement.
Industry relevance. The pipeline is model-agnostic in structure: the authors ran all experiments with gemini-3.1-flash as the understanding MLLM and gemini-3.1-flash-image as the image generator, meaning the method can be layered on existing commercial multimodal APIs. That makes it a candidate for design-assistant products, generative art tools, and advertising or branding workflows, where the barrier to adoption is orchestration logic rather than new model training.
Future Directions
- Quantitative comparison against closed-source baselines. The current baseline comparison is qualitative; extending the N = 50 evaluation protocol to a head-to-head study against Gemini, GPT, and Claude would test whether the reported advantages hold under controlled measurement.
- Human-in-the-loop evaluation. The authors highlight that preserving the morphology of a designer-selected base form is crucial for human-in-the-loop workflows, but no study with designers steering the base image is reported.
- Generalization beyond the reported topics. Results are shown on a small set of abstract topics and a typography series; whether the staged framework scales to a broad, systematic topic set is not established.
- Robustness of the intermediate stages. The ablation shows that an underspecified base prompt propagates a semantic error through the entire pipeline, which raises the open question of how to detect and correct failures at the intermediate stages rather than at the final image.
Target Audience
This paper is most useful to researchers and practitioners working on multimodal agents and controllable image generation, especially those interested in design-oriented tasks that require joint spatial and semantic control rather than simple prompt-to-image mapping. It also serves graphic designers, creative technologists, and product teams evaluating MLLM-based agent pipelines for creative tooling, as well as students studying the gap between visual understanding and visual generation in multimodal models. Readers seeking a quantitative benchmark study or a trained generative architecture will find the paper's evaluation qualitative and ablation-focused rather than benchmark-driven.
Authors’ abstract
Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbf{F}orm \textbf{a}nd \textbf{V}oid \textbf{A}gent (\textbf{FaV-A}), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.