Skip to content
AI.info

Research

DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control

Overview Research area: Computer vision, specifically controllable text-to-image generation using diffusion models, with a focus on multi-instance generation (generating several distinct objects or pe

DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control
arXiv
2602.18282
Published
2026-02-20
Authors
Shiyan Du, Conghan Yue, Xinyu Cheng, Dongyu Zhang

AI summary

Overview

Research area: Computer vision, specifically controllable text-to-image generation using diffusion models, with a focus on multi-instance generation (generating several distinct objects or people at user-specified locations).

Technical level: Advanced. The paper assumes familiarity with diffusion models, cross-attention and self-attention mechanisms, text encoders, and region-level (bounding-box) conditioning.

Scope: The paper introduces DEIG, a framework plus a new benchmark (DEIG-Bench) for generating multi-instance images that honor fine-grained, multi-attribute per-instance descriptions such as color, material, and texture.

What This Paper Is About

Existing multi-instance generation methods can place objects at specified bounding boxes and bind simple attributes like "a red car," but they break down when descriptions become rich and compositional, for example a person wearing multiple colored garments or a pillow described by color, material, and texture at once. The authors trace this to two causes: methods focus on stopping semantic leakage between instances while neglecting deeper semantic comprehension, and training data uses coarse templates rather than detailed instance-level captions. DEIG addresses both by adding instance-aware semantic extraction and masked attention fusion, and by training on a new VLM-captioned dataset.

Key Contributions

  1. DEIG framework: A pipeline that enhances instance-level detail representation and semantic understanding, built around two new modules — the Instance Detail Extractor (IDE) and the Detail Fusion Module (DFM) — to overcome the limits of existing methods on rich, fine-grained region descriptions.

  2. DEIG-Bench: A new evaluation suite for multi-attribute, multi-instance generation, designed to fill the gap of missing benchmarks for fine-grained semantic prompts. It contains region-level annotations and multi-attribute prompts for both humans (levels C1–C3, based on color combinations across wearable regions) and objects (levels L1–L4, adding material and texture).

  3. A detail-enriched instance caption dataset: Built from MS-COCO using Qwen2.5-VL to generate detailed, context-aware captions averaging 20–30 words per instance, replacing template-based annotations.

  4. Extensive experiments: Evaluations on DEIG-Bench, MIG-Bench, and InstDiff-Bench showing DEIG outperforms previous methods, plus ablations and a plug-and-play demonstration on a community diffusion backbone.

Main Findings

  • DEIG leads on DEIG-Bench human attributes: Using Qwen2.5-VL as evaluator, DEIG scores Multi-Attribute Accuracy (MAA) of 0.82 / 0.74 / 0.69 across complexity levels C1–C3 (average 0.75) and 0.67 / 0.41 / 0.38 / 0.27 across object levels L1–L4 (average 0.44). With InternVL3 the same settings give human scores of 0.86 / 0.82 / 0.81 (average 0.83) and object scores of 0.79 / 0.58 / 0.50 / 0.44 (average 0.58).

  • Baselines trail substantially on human instances: Under Qwen2.5-VL, ROICtrl reaches a human average of 0.31, InstanceDiffusion 0.25, MIGC 0.22, and GLIGEN 0.10. The paper attributes the gap to stronger compositional generalization from the detail-aware extraction and fusion modules.

  • Color is easier than material and texture: Object-level gains are larger on color and more modest on material and texture, which the authors explain by color correlating more directly with RGB space while material and texture require more abstract semantic understanding.

  • Strong results on MIG-Bench: DEIG achieves the best average Instance Success Rate (72.25) versus MIGC (65.84), ROICtrl (63.25), InstanceDiffusion (58.63), and GLIGEN (29.91), and the best average mIoU (62.64) versus MIGC (56.44), ROICtrl (55.27), InstanceDiffusion (53.06), and GLIGEN (27.03).

  • Strong attribute alignment but slightly lower spatial precision on InstDiff-Bench: DEIG reports Accuracy for color 58.8 and CLIP color 0.258, Accuracy for texture 26.1 and CLIP texture 0.228, with AP 0.34 and AP50 0.57. InstanceDiffusion reports the higher AP of 0.40 with the same AP50 of 0.57; the authors attribute DEIG's slightly lower spatial alignment to instance-masked attention limiting interactions in crowded regions.

  • Caption supervision matters most in ablation: Removing the detail-enriched captions drops MAA-human to 0.51 and MAA-obj to 0.35 (from 0.75 and 0.44 with all components). Removing IDE gives 0.31 / 0.29, and removing DFM gives 0.70 / 0.41. mIoU falls to 0.73 and 0.70 respectively in the first two settings. Object accuracy degrades less than human accuracy, indicating human-centric generation is more sensitive to loss of fine-grained control.

  • Aggregated semantic dimension S has a precision-cost trade-off: Increasing S improves MAA for both human and object instances, with gains saturating around S = 16; beyond that, performance plateaus or slightly declines due to overfitting while GPU memory under FP16 rises steadily. The authors recommend S = 16 ~ 32 as a balance.

  • Plug-and-play compatibility: DEIG adapts to a community diffusion backbone without retraining, preserving spatial layout and fine-grained generation quality.

Methodology in Plain English

The system takes a global prompt plus a list of instances, where each instance is a bounding box paired with a detailed text description. The goal is one image that satisfies both the spatial and the semantic constraints.

The original text embeddings from a frozen Flan-T5-XL encoder are high-dimensional and expensive to use directly. The Instance Detail Extractor (IDE) compresses them into a small set of learnable queries per instance. Each IDE layer applies timestep conditioning through a lightweight TimeMLP, adaptive layer normalization (AdaLN), a self-attention block for intra-instance dependencies, and a cross-attention block that aligns the queries with the frozen encoder's text features, followed by a residual feed-forward network. The number of queries is the "Aggregated Semantic Dimension" S, which acts as a bottleneck. Six such layers (N = 6) are stacked, producing "Aggregated Semantic Embeddings." Visualizing these, the authors find each dimension attends to specific fine-grained attributes, and together they form a complete representation of the instance.

The Detail Fusion Module (DFM) injects those embeddings into generation. First, Grounding Embeddings Broadcast aligns spatial cues with the semantic dimension: each instance's bounding-box coordinate is Fourier-encoded and broadcast across all S dimensions, then combined with a learnable null embedding (chosen by a binary mask when spatial information is absent) and passed through an MLP to produce a fused spatial-semantic embedding.

Second, Instance-based Masked Attention prevents attribute leakage. The self- and cross-attention layers of the UNet stay frozen, and a gated self-attention module is inserted between them. The attention map naturally splits into four sub-regions, and the authors define masking rules: visual embeddings attend to each other freely (masking them hurts fidelity); instance embeddings attend only to visual embeddings of the same instance and vice versa, with cross-instance scores set to negative infinity; and instance embeddings attend only within the same semantic group, again masking cross-group interactions. The masked output feeds a gated residual update with learnable scalars controlling update strength.

Training data comes from MS-COCO, recaptioned with Qwen2.5-VL. Grayscale and low-fidelity images are removed by VLM assessment, and a two-stage verification keeps image–caption pairs above a predefined CLIP score threshold, followed by human verification on a random subset of 500 pairs. Models are trained at 512×512 resolution, initialized from a pre-trained GLIGEN checkpoint based on Stable Diffusion v1.4, for 800k iterations on 8 NVIDIA RTX 3090 GPUs, with AdamW at a constant learning rate of 1e-4, a linear warm-up over the first 10k iterations, batch size 4 with gradient accumulation of 4 (effective batch size 128), S = 16, and N = 6 IDE layers.

DEIG-Bench itself is built from 400 filtered MS-COCO validation images, each containing 3–10 visible instances; instances were retained only if their relative area was between 10% and 60% of the image. Its attribute vocabulary includes 13 common colors, 8 canonical materials, and 4 textures or patterns. Object descriptions use a probabilistic composition: 30% color only, 25% color plus material, 25% color plus texture, and 20% all three. Evaluation combines mIoU computed with Grounding-DINO for spatial alignment, two VLMs in a question-answering setup for semantics, and the newly introduced Multi-Attribute Accuracy (MAA), the ratio of instances correctly identified by the VLM to total instances.

Why This Matters

Impact on research. The paper reframes multi-instance generation as a problem of semantic depth rather than only leakage prevention, and it argues that annotation granularity is a first-class bottleneck — template captions limit what a model can learn. It also supplies a benchmark targeting two gaps in prior evaluation: underrepresentation of human instances and reliance on single-attribute prompts. Because DEIG is a plug-and-play module, the approach can be layered onto existing diffusion pipelines rather than requiring a new architecture.

Real-world applications (as described in the paper):

  • Fashion synthesis and artistic creation, where garments and regions need distinct attributes.
  • Animation and design, where content production benefits from efficiency and consistency.
  • Advertising, where specific products must appear at specific positions with specific colors, materials, and textures.
  • Education, where lowering the technical barrier lets a wider range of users produce high-quality visuals.

Industry relevance. The authors frame the work as democratizing creative content production while also noting risks: highly controllable synthesis could be misused for misleading content, automated tools may disrupt traditional creative workflows and create economic challenges for professionals, and dataset biases may be amplified in outputs. They call for technical safeguards such as content filtering and usage guidelines, plus responsible deployment practices. Training requirements (800k iterations on 8 RTX 3090 GPUs) are modest compared with DiT-based methods, which the paper notes have high computational overhead that makes them impractical on consumer-grade GPUs.

Future Directions

  • Overlapping and dense scenes. The model can fail to disentangle adjacent instances when objects overlap heavily, as seen on benchmarks like InstDiff-Bench, producing incomplete or unstable generations. The authors state that future work will focus on improving the handling of overlapping regions.

  • Small objects. Preserving fine-grained details for small objects remains difficult because limited spatial resolution hinders accurate attribute depiction.

  • Material and texture understanding. Because gains on material and texture are more modest than on color, closing this gap is a natural extension of the work on abstract semantic attributes.

  • Broader evaluation and integration. The paper presents DEIG-Bench, MIG-Bench, and InstDiff-Bench results and demonstrates plug-and-play adaptation to a community diffusion backbone; scaling this compatibility and the benchmark's coverage of unconstrained, free-form generation are open threads.

Target Audience

Researchers and graduate students in generative vision working on controllable and compositional text-to-image synthesis; engineers building layout-conditioned image generation systems for fashion, design, or advertising; and benchmark developers interested in region-level, multi-attribute evaluation protocols for both human and object instances.

Authors’ abstract

Multi-Instance Generation has advanced significantly in spatial placement and attribute binding. However, existing approaches still face challenges in fine-grained semantic understanding, particularly when dealing with complex textual descriptions. To overcome these limitations, we propose DEIG, a novel framework for fine-grained and controllable multi-instance generation. DEIG integrates an Instance Detail Extractor (IDE) that transforms text encoder embeddings into compact, instance-aware representations, and a Detail Fusion Module (DFM) that applies instance-based masked attention to prevent attribute leakage across instances. These components enable DEIG to generate visually coherent multi-instance scenes that precisely match rich, localized textual descriptions. To support fine-grained supervision, we construct a high-quality dataset with detailed, compositional instance captions generated by VLMs. We also introduce DEIG-Bench, a new benchmark with region-level annotations and multi-attribute prompts for both humans and objects. Experiments demonstrate that DEIG consistently outperforms existing approaches across multiple benchmarks in spatial consistency, semantic accuracy, and compositional generalization. Moreover, DEIG functions as a plug-and-play module, making it easily integrable into standard diffusion-based pipelines.

Read the original paper