Research
Compositional Image Synthesis with Inference-Time Scaling
Compositional Image Synthesis with Inference-Time Scaling Overview Research area: Computer vision / generative modeling — specifically text-to-image (T2I) diffusion, compositional (layout-faithful) im

- arXiv
- 2510.24133
- Published
- 2025-10-28
- Authors
- Minsuk Ji, Sanghyeok Lee, Namhyuk Ahn
AI summary
Compositional Image Synthesis with Inference-Time ScalingOverview
Research area: Computer vision / generative modeling — specifically text-to-image (T2I) diffusion, compositional (layout-faithful) image synthesis, and inference-time scaling with vision-language model judging.
Technical level: Intermediate. The paper assumes familiarity with diffusion models, classifier-free guidance, CLIP similarity, layout-grounding conditioning, and VLMs used as judges.
Scope: The paper introduces ReFocus, a training-free framework that combines LLM-generated layouts, layout-grounded diffusion drafting, and object-centric iterative self-refinement with re-ranking to improve compositional faithfulness of text-to-image generation on the GenEval benchmark (arXiv:2510.24133v2 [cs.CV], 27 Mar 2026; authors Minsuk Ji*, Sanghyeok Lee*, Namhyuk Ahn, Inha University; * indicates equal contribution).
What This Paper Is About
Modern text-to-image diffusion models produce highly realistic images but still fail at compositionality — rendering the correct object counts, attributes, and spatial relations (for example, "a photo of four giraffes" often yields the wrong number of animals, and "a photo of a chair left of a zebra" can produce spatial inconsistency). Prior fixes either force users to hand-draw layouts (e.g., bounding boxes) or rely on scene-level Best-of-N selection that never fixes fine-grained, object-level errors. The goal is a user-friendly, training-free framework that imposes explicit compositional structure while preserving aesthetic quality.
Key Contributions
-
A training-free end-to-end framework (ReFocus) that unifies object-centric layout grounding with self-refine-based inference-time scaling, requiring no feedback collection or alignment tuning. Table 1 positions ReFocus as the only listed method combining all four properties — object-centric, inference-time scaling, self-refine, and training-free — in contrast to SD1.5/SDXL/FLUX (training-free only), GLIGEN/ControlNet (object-centric only, not training-free), Best-of-N/Z-Sampling (inference-time scaling only), and Reflect-DiT (which requires training).
-
Automatic LLM-based layout generation. Instead of manual layout annotation, an LLM (ChatGPT-4o in the implementation) parses an unstructured prompt into an explicit layout of object labels and normalized bounding boxes, removing the cumbersome user burden of supplying boxes.
-
A hybrid scene-plus-object re-ranking score. Rather than scene-level CLIP similarity alone, the framework combines a scene-level CLIP score with an object-level score computed by cropping object regions from each draft per the generated layout and averaging CLIP similarity over them, weighted by a hyperparameter λ.
-
Iterative refinement coupled with re-ranking. A lightweight refinement model generates variants from the top-K re-ranked candidates at low denoising strength, and the re-rank/refine loop repeats — improving realism without losing the geometric arrangement established by layout grounding.
Main Findings
-
Highest average GenEval score: ReFocus (N=4) reaches 0.84 average on GenEval, above Reflect-DiT + Best-of-N (N=20) at 0.81, Sana-1.0-1.6B + Best-of-N (N=20) at 0.75, SD3 at 0.74, FLUX.1-dev at 0.68, DALL-E 3 at 0.67, SDXL + GLIGEN at 0.65, SDXL + Best-of-N (N=4) at 0.61, SDXL + Z-Sampling at 0.57, SDXL at 0.55, SD1.5 + Best-of-N (N=4) at 0.51, and SD1.5 at 0.43.
-
Large gains on the hardest categories: ReFocus scores 0.81 on Position and 0.82 on Counting. Compared with the SDXL baseline, the paper reports improvements of +0.66 on position and +0.43 on counting. It also yields higher position accuracy than GLIGEN, which explicitly conditions on layouts.
-
Fewer samples, stronger results: Against Reflect-DiT (which also combines inference-time scaling with self-revision), ReFocus is training-free — no feedback collection or alignment tuning — and achieves better performance with fewer inference samples (N=4 versus N=20).
-
Scaling behavior holds: Figure 2 shows the average GenEval score as the number of samples per prompt (N) increases; the ReFocus curve consistently stays above competing methods and the gap remains as N grows, suggesting a scalable solution rather than a one-off heuristic.
-
Ablation (Table 3): Phase 1 (layout grounding only) achieves GenEval 0.78 but limited realism (HPS v2 average 26.74). Adding inference-time scaling (Best-of-N) gives 0.80 GenEval and 27.17 HPS v2 — modest gains. Adding refinement (Round 1) gives 0.80 GenEval and 28.29 HPS v2, and Round 2 reaches 0.84 GenEval and 28.32 HPS v2. For reference, SD1.5 scores 0.43 GenEval / 27.12 HPS v2 and SD2.0 scores 0.51 / 27.17.
-
Qualitative behavior: The paper reports that sample prompts such as "a purple elephant and a brown sports ball" and "four traffic signs" show accurate object counts, faithful colors, and correct relative positions, whereas prior methods often miss objects or distort spatial relations. Figure 4 illustrates that layout grounding establishes geometry, re-ranking selects the most prompt-aligned candidate, and refinement sharpens texture, lighting, and boundaries while preserving geometry.
-
Where the difficulty lies: The authors note that because the backbone diffusion model (SDXL) has trouble with overlapping objects, the initial drafts can be of poor quality in such cases, motivating the refinement phase.
Methodology in Plain English
The framework runs in three phases and requires no additional model training.
Phase 1 — Layout generation. An LLM reads the text prompt and outputs a layout: a set of object labels, each paired with a bounding box in normalized coordinates (x_min, y_min, x_max, y_max) in [0,1]. Unlike prior layout-grounding methods, the user does not have to draw the boxes.
Phase 1 detail — margin strategy. Boxes are shrunk by an adaptive margin δ in [0.02, 0.04]. The lower value is calibrated to the Latent Diffusion architecture: given the 1/8 downsampling factor (512×512 → 64×64), a margin of 0.02 corresponds to roughly 1.28 pixels in latent space, acting as a boundary against "concept bleeding" between adjacent objects. Rather than a rigid non-overlap constraint, the LLM distinguishes independent from interacting objects: margins are relaxed for objects with explicit depth dependencies (e.g., "behind") to allow natural occlusion, and tightened for spatially distinct objects.
Phase 2 — Layout-grounded drafting. A layout-conditioned diffusion model G samples N draft images from independent standard Gaussian noise, each conditioned on both the prompt and the layout. These drafts impose a coarse compositional structure from the start; unlike existing layout-grounding methods, they are not treated as the final output but as the basis for refinement.
Phase 3 — Iterative self-refinement. Drafts are re-ranked with a hybrid score: S = λ·S_scene + (1−λ)·S_object, where S_scene is standard CLIP similarity between image and prompt, and S_object averages CLIP similarity over object crops taken according to the layout. The top-K candidates are then passed to a lightweight refinement model that applies independent noise and partial denoising at a low denoising strength (α_refine ≪ 1), producing M variants. Re-ranking and refinement repeat in a loop, so errors such as missing objects or implausible details can be corrected while realism improves.
Implementation specifics: MIGC serves as the layout-conditioned diffusion model G in Phase 2; SDXL-Turbo is the refinement model in Phase 3, chosen for fast inference to minimize latency in the iterative loop; ChatGPT-4o performs layout parsing in Phase 1. Phase 2 uses 50 sampling steps and a classifier-free guidance scale of 7.5. Phase 3 uses a single refinement step with guidance fixed at 0.0 and denoising strength 0.5.
Evaluation: GenEval measures object-level compositional accuracy; HPS v2.1 measures visual quality and human preference. Baselines include representative diffusion text-to-image models, a layout-grounding model, inference-time scaling methods, and a feedback-based scaling method.
Why This Matters
Impact on research. The paper argues that explicit layout grounding plus object-centric refinement is a scalable solution rather than a one-off heuristic, since the ReFocus curve stays above competing methods as N grows. It also shows that object-level judging can substitute for the expensive reflection-tuning that methods like Reflect-DiT require, and that a hybrid scene-plus-object preference score is a stronger selection signal than scene-level CLIP alone — a design pattern transferable to other inference-time scaling pipelines.
Real-world applications (derived from the paper's demonstrated capabilities; the paper does not itself enumerate deployment scenarios):
- Design and marketing asset creation — prompts demanding exact object counts, colors, and left/right relations (e.g., "four traffic signs") that generic T2I models routinely get wrong.
- E-commerce and catalog imagery — generating scenes with a specified number of products and prescribed placement without a human drawing bounding boxes.
- Storyboarding and illustration — relational prompts such as "a chair left of a zebra" where spatial consistency is the requirement, not an afterthought.
- Accessible content generation — automatic layout inference from plain text means non-expert users get layout-grounding benefits without layout-authoring tools.
Industry relevance. Because the framework is training-free and modular (an LLM parser, an existing layout-conditioned diffusion backbone, and a fast refinement model such as SDXL-Turbo), it can be layered on top of existing generative pipelines without collecting feedback data or running alignment training. The code is publicly available at https://minsuk-ji.github.io/ReFocus/. The work was supported by the Institute of Information & communications Technology Planning & Evaluation (IITP) under the Leading Generative AI Human Resources Development grant (IITP-2026-RS-2024-00360227), funded by the Korea government (MSIT).
Future Directions
-
Reducing reliance on model components. The pipeline currently depends on several external models (ChatGPT-4o for layout parsing, MIGC for layout-conditioned drafting, SDXL-Turbo for refinement). Whether cheaper or open alternatives preserve the reported GenEval and HPS v2 numbers is an open question not addressed in the paper.
-
Overlapping and occluded objects. The paper explicitly notes that the backbone diffusion model struggles with overlapping objects, producing poor-quality drafts in those cases; better handling of heavy occlusion and interaction is a natural next step.
-
Margins, thresholds, and other hyperparameters. The margin range δ ∈ [0.02, 0.04], the balance weight λ, the number of retained top-K candidates, and the refinement denoising strength (0.5) are set by the authors. The paper does not report a sensitivity analysis over these values, leaving their generalization across backbones unclear.
-
Scaling behavior beyond the tested regime. Figure 2 studies GenEval score as N increases; how far the reported advantage over methods like Reflect-DiT + Best-of-N extends at much larger N, and its cost/quality trade-off, is not reported.
Target Audience
Researchers and practitioners in generative computer vision who work on text-to-image diffusion, compositional or layout-conditioned synthesis, and inference-time scaling; engineers building image-generation products who need prompt-faithful outputs without training custom models; and readers interested in LLM-as-layout-parser plus VLM-as-judge pipelines as a training-free alternative to reflection-tuned generators.
Authors’ abstract
Despite their impressive realism, modern text-to-image models still struggle with compositionality, often failing to render accurate object counts, attributes, and spatial relations. To address this challenge, we present a training-free framework that combines an object-centric approach with self-refinement to improve layout faithfulness while preserving aesthetic quality. Specifically, we leverage large language models (LLMs) to synthesize explicit layouts from input prompts, and we inject these layouts into the image generation process, where a object-centric vision-language model (VLM) judge reranks multiple candidates to select the most prompt-aligned outcome iteratively. By unifying explicit layout-grounding with self-refine-based inference-time scaling, our framework achieves stronger scene alignment with prompts compared to recent text-to-image models. The code are available at https://github.com/gcl-inha/ReFocus.