Skip to content
AI.info

Research

OverLay++: Dense-Overlap Layout-to-Image Generation Dataset

Overview Research area: Controllable image generation, specifically Layout-to-Image (L2I) synthesis, with a focus on training-data construction rather than model architecture. Technical level: Interme

OverLay++: Dense-Overlap Layout-to-Image Generation Dataset
arXiv
2610.09071
Published
2026-10-06
Authors
Shivansh Aggarwal, Shresth Grover, Divyansh Srivastava, Haiyang Xu, Bingnan Li, Xiang Zhang, Ethan J. Armand, Chuan Li, Jianwen Xie, Zhuowen Tu

AI summary

Overview

Research area: Controllable image generation, specifically Layout-to-Image (L2I) synthesis, with a focus on training-data construction rather than model architecture.

Technical level: Intermediate. The paper is written in accessible prose, but it assumes familiarity with diffusion backbones, bounding-box conditioning, and evaluation metrics such as FID, IS, mIoU, and CLIP scores.

Scope: The paper introduces OverLay++ (arXiv:2610.09071v1, cs.CV, 06 Oct 2026), a roughly 500K-image Layout-to-Image dataset built from an automated four-stage pipeline to supply denser object annotations, more overlapping objects, and richer multi-level captions than prior L2I datasets, and it benchmarks state-of-the-art L2I methods trained on it.

What This Paper Is About

Layout-to-Image models condition a text-to-image diffusion model on a scene layout — a set of bounding boxes, each paired with a local caption — plus a global caption. The authors argue that these models fail on cluttered scenes with many overlapping objects not because of architecture, but because existing training datasets contain few dense, overlapping layouts and their local annotations are mostly short class names or phrases. The goal of the paper is to fix the data bottleneck by generating a large dataset whose layouts are explicitly dense and overlap-aware, and whose captions describe object appearance at both short and long granularity.

Key Contributions

  1. The OverLay++ dataset. Approximately 500K images with an average of 6.59 objects per image, exceeding existing datasets by 1.67 times in annotation density, with short and long captions at both the global and per-object levels. The paper reports that even its short captions are 6 times longer than those in existing datasets.

  2. A four-stage automated curation pipeline. The pipeline filters images by aesthetics and resolution, detects objects with an MLLM while explicitly asking for overlapping bounding boxes, applies per-object grounded validation to remove hallucinations, and produces hierarchical short/long global and local captions.

  3. Benchmarking of state-of-the-art L2I methods on OverLay++. EliGen-SD3 and SiamLayout-SD3 trained or fine-tuned on OverLay++ are evaluated on three benchmarks (LayoutSAM, OverLayBench, DenseLayout), with consistent improvements over their corresponding baselines and faster convergence.

  4. A controlled study of caption granularity. An ablation over short/long local and global captions shows that short local captions with long global captions give the best region-level alignment, with the authors hypothesizing that the 77-token CLIP path truncates long local captions.

Main Findings

  • Dataset statistics (Table 1). OverLay++ reports 500K images, 6.59 objects per image, mean nearest-neighbor IoU 0.31, mean nearest-neighbor distance 0.17, local caption lengths of 15.89 (short) and 89.70 (long) words, and global caption lengths of 19.98 (short) and 52.24 (long) words. By comparison, LayoutSAM reports 2.66M images at 3.93 objects, EliGen reports 496K at 2.56 objects, and GRIT reports 20.5M at 1.78 objects; prior datasets range from 1.78 to 3.93 objects per image.

  • LayoutSAM benchmark (Table 2, LayoutSAM-Eval, 5,000 layouts). EliGen-SD3 (Ours) reaches 95.45% Spatial (+2.59%), 90.03% Color (+8.52%), 93.14% Texture (+7.39%), and 92.81% Shape (+7.48%), with FID 19.61 (−22.09%) and IS 19.54 (+1.72%). SiamLayout-SD3 (Ours) reaches 93.30 Spatial (+0.68%), 75.99 Color (+2.07%), 78.77 Texture (+2.02%), 78.41 Shape (+3.27%), FID 17.96 (−5.97%), and IS 21.60; the paper describes this as the best FID and second-best IS among all compared methods.

  • OverLayBench benchmark (Table 3). Across Simple, Regular, and Complex splits, EliGen-SD3 (Ours) achieves the best mIoU, O-mIoU, and SR_E on Regular (62.40, 36.66, 89.45%) and Complex (57.90, 30.99, 83.30%), and the best SR_E and SR_R on Simple (91.83% and 91.06%). Gains over the EliGen-SD3 baseline reach +14.59% in mIoU and +34.92% in O-mIoU on the Complex split. Not every metric improves: CLIP_G drops slightly for EliGen-SD3 (Ours) on all three splits (e.g., 35.46 vs 36.49 on Complex), and FID changes are marginal on Simple (−0.03%) and Regular (−0.12%).

  • DenseLayout benchmark (Table 4). EliGen-SD3 (Ours) improves mIoU from 45.79 to 46.92 and gains 7.86% in Color, 7.33% in Texture, and 7.70% in Shape. SiamLayout-SD3 (Ours) improves on all four metrics, with gains of 10.04% mIoU, 12.49% Color, 11.59% Texture, and 11.83% Shape over its baseline.

  • Faster convergence. The convergence analysis (Figure 4) compares mIoU and O-mIoU against training steps for the Simple, Regular, and Complex settings; the model trained on OverLay++ reaches higher scores at all difficulty levels and converges faster than the same method trained on the EliGen dataset.

  • Caption granularity matters (Table 5). Short local captions paired with long global captions give the best overall results for EliGen-SD3 on OverLayBench, while long local captions perform substantially worse (for example, Complex mIoU 57.90 for Short-Long versus 47.64 for Long-Long and 44.23 for Long-Short). The authors attribute this to the 77-token limit of the CLIP text encoder used for local prompts, noting long local captions average 89.70 words and that this is a hypothesis rather than a direct measurement of truncation.

  • Long-context text encoder (Appendix A.2, Tables 6 and 7). Fine-tuning the released EliGen-FLUX model for 5K steps with batch size 8 on OverLay++ using long local captions improves nearly every metric across all OverLayBench splits, with SR_E improving by up to 11.54% on the Complex split, and improves all region-wise metrics plus FID and IS on LayoutSAM-Eval (FID 15.98, a 44.88% improvement over EliGen-FLUX).

  • Long-CLIP alignment (Appendix A.3, Table 8). OverLayBench global prompts average 79.66, 86.81, and 82.91 words on the Simple, Regular, and Complex splits, exceeding CLIP's 77-token limit. Re-scored with Long-CLIP (up to 248 tokens), the OverLay++-trained EliGen-SD3 model has slightly higher scores than its baseline on all three splits (0.2573 vs 0.2564, 0.2559 vs 0.2556, 0.2570 vs 0.2563).

  • Not reported in the supplied content. The ablation of object count and local caption length (Table 9) is cut off mid-sentence in the available text; the paper states only that these models use the SD3 backbone and are trained for 2K steps on OverLay++ subsets with matched mean IoU.

Methodology in Plain English

The authors start from the LAION-Aesthetics V2 4.75 subset, which contains 956M images, and run four automated stages.

  1. Image filtering. They keep only images with a LAION aesthetic score of at least 5.98 and a minimum spatial dimension of at least 1024 pixels, then center-crop to 1024 × 1024. This leaves 850,900 images.

  2. Object detection. They prompt Qwen3-VL-32B to detect all visually salient objects and return a category label, a bounding box normalized to a 1024-pixel coordinate space, and a detailed caption describing the object. The prompt explicitly instructs the model that each bounding box must overlap at least one other bounding box and must not be very small (minimum 51 pixels in each dimension). Images with fewer than five valid detections are discarded.

  3. Scene filtering. Each candidate detection is cropped and re-checked with Qwen3-VL-32B using a yes/no prompt asking whether the described object is present in the crop; negative detections are removed, and images with fewer than five valid detections remaining are discarded. This leaves 499,249 images.

  4. Captioning. Qwen3-VL-32B produces both scene-level and object-level captions in long and short forms. Short captions are prompted to stay under 20 words; the word limits are prompt instructions rather than post-generation truncation.

To evaluate, the authors train EliGen from public code on both the EliGen dataset and OverLay++ starting from an SD3 backbone, for 10K steps with per-GPU batch size 12 on 8 × A6000 GPUs at a learning rate of 3 × 10⁻⁴, using short local captions and long global captions. Because SiamLayout's training code is not public, they instead fine-tune the released SiamLayout model on OverLay++ for 60K steps with per-GPU batch size 2, gradient accumulation 2, and learning rate 5 × 10⁻⁵ on 8 × A6000 GPUs, comparing against SiamLayout-SD3 numbers reported in the original paper. All runs use AdamW with weight decay 0.05, β₁ = 0.9, β₂ = 0.95, a constant learning rate, and a classifier-free guidance scale of 7.5. Evaluation covers LayoutSAM-Eval, OverLayBench (Simple, Regular, Complex), and DenseLayout.

Why This Matters

Research impact. The paper reframes dense-scene L2I failures as a data problem rather than a modeling problem. It supplies both a dataset and a reproducible generation recipe, and it demonstrates that the same methods improve simply by swapping the supervision data, which gives the community a clear baseline to build on. Its caption-granularity and long-context analyses also connect text-encoder token limits to measurable generation quality.

Real-world applications:

  • Design and advertising. Generating product mockups or ad creatives from a specified layout, where objects must appear in exact positions and retain specified colors, textures, and shapes.
  • E-commerce and catalog imagery. Producing lifestyle scenes in which several products appear together with correct spatial arrangement and overlap.
  • Content creation and media production. Turning storyboard-style layouts into images for previsualization, where dense, interacting characters and props are common.
  • Synthetic data generation. Creating dense, accurately annotated scenes to train or augment downstream perception models.

Industry relevance. Dense, overlap-heavy scenes are exactly where controllable generation is most commercially useful and most brittle. A dataset that improves region-level spatial and attribute alignment reduces retry loops and manual correction in production pipelines, and the finding that models converge faster on OverLay++ lowers the compute cost of adapting L2I models to new domains — relevant to any team training or fine-tuning diffusion models with layout control.

Future Directions

  • Test long local captions with longer-context encoders. The degraded performance of long local captions is attributed to CLIP's 77-token limit; the authors note that stronger local text encoders or longer context windows may unlock additional gains from long local captions, and the FLUX results in the appendix are presented as supporting evidence for that hypothesis rather than proof.
  • Close the remaining density gaps. The limitation section acknowledges that the dataset may underrepresent unusual object interactions, transparent objects, and rare object categories.
  • Reduce inherited model bias and annotation noise. Because the pipeline relies on MLLMs for captioning and grounding, the dataset may inherit model biases and annotation errors; better validation or human-in-the-loop correction is an open problem, and the over/under-representation of rare categories is not resolved.
  • Study misuse of L2I models. The authors call for further research on misuse risks to support safe deployment of controllable image generation systems.

Target Audience

Researchers and engineers working on controllable image generation, diffusion model conditioning, and layout-to-image synthesis, as well as dataset builders who need a concrete pipeline for generating dense, richly captioned ground-truth annotations. It is also useful to practitioners fine-tuning L2I models who want to know which caption granularity to use (short local, long global) and what quality gains to expect from better supervision data. The evaluation results and ablation on caption granularity assume some familiarity with diffusion backbones and standard generation metrics.

Authors’ abstract

Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex object interactions. To address this gap, we introduce OverLay++, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by 1.67 times in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than in current datasets. Our dataset generation pipeline is simple and produces dense, overlapping object annotations with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on the OverLay++ dataset show consistent improvement and faster convergence, demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.

Read the original paper