Skip to content
AI.info

Research

Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment

Overview Research area: Text-to-image (T2I) generation, specifically text-image alignment in diffusion and flow-matching models, and the use of negative prompting within classifier-free guidance (CFG)

Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
arXiv
2512.07702
Published
2025-12-08
Authors
Sangha Park, Eunji Kim, Yeongtak Oh, Jooyoung Choi, Sungroh Yoon

AI summary

Overview

Research area: Text-to-image (T2I) generation, specifically text-image alignment in diffusion and flow-matching models, and the use of negative prompting within classifier-free guidance (CFG).

Technical level: Intermediate. Readers benefit from familiarity with diffusion/flow-matching sampling, classifier-free guidance, and transformer cross-attention, though the paper explains its attention analysis from first principles.

Scope: The paper proposes NPC (Negative Prompting for Image Correction), a fully automated pipeline that generates and ranks candidate negative prompts — both targeted (tied to a specific alignment failure) and untargeted (incidental content present in the generated image) — to improve text-image alignment without manual prompt engineering.

What This Paper Is About

Text-to-image models often fail to follow prompts that involve multiple objects, attributes, spatial relations, counts, or surreal instructions, even though they produce photorealistic images. Prior work has largely focused on positive guidance — strengthening conditioning on what the prompt asks for — or on layout-centric and targeted correction pipelines that suppress content directly contradicting the prompt. This paper asks whether telling the model what not to generate, including negatives that are not directly tied to the alignment error, can improve alignment, and it builds an automated pipeline to find and select such negatives.

Key Contributions

  1. Broadened scope of negative prompting with an attention-based explanation. The authors analyze image-text cross-attention in transformer-based diffusion/flow models and show that both targeted negatives (directly tied to the alignment error) and untargeted negatives (tokens unrelated to the prompt but present in the generated image) increase attention mass on the positive prompt's salient tokens.

  2. An automated negative-prompt generation pipeline without manual prompting. A verifier–captioner–proposer framework produces a set of candidate negative prompts: the verifier returns a failure reason (a source of targeted negatives), the captioner describes the generated image (a source of untargeted negatives), and the proposer distills both into short candidates.

  3. A text-space scorer that avoids extra image synthesis. The proposed salient score ranks candidates using only the text encoder, so that most candidates can be evaluated without generating images.

  4. State-of-the-art results on two challenging benchmarks. NPC achieves 0.571 overall on GenEval++ versus 0.371 for the strongest baseline (Bagel), and the best overall performance on Imagine-Bench, with additional results on T2I-CompBench and DPG-Bench reported in the appendix.

Main Findings

  • GenEval++ overall score: NPC reaches 0.571, compared with 0.371 for Bagel (the strongest baseline), a difference of 0.200, or roughly a 54% relative gain. NPC's per-category scores are 0.550 (Color), 0.675 (Count), 0.350 (Color/Count), 0.525 (Color/Pos), 0.550 (Pos/Count), 0.725 (Pos/Size), and 0.625 (Multi-Count). The paper reports that NPC ranks second best on Color and Color/Pos while leading on Count, Pos/Count, Pos/Size, and Multi-Count.

  • Imagine-Bench overall score: NPC achieves the best overall score of 6.80, ahead of T2I-R1 at 6.78 and Bagel at 6.20. NPC ranks first in every category except Spatiotemporal, where T2I-R1 leads (7.70 vs. NPC's 6.76). NPC's other category scores are 6.31 (Attribute shift), 7.57 (Hybridization), and 6.52 (Multi-Object).

  • Attention analysis of negative prompts: For the prompt "A photo of a green fire hydrant, a black elephant, and a brown tv," the BASE configuration (no negative prompt) attains a salient-attention score of 0.069. Targeted negative guidance raises it to 0.096 (a 39.13% increase) and an untargeted negative raises it to 0.075 (an 8.70% increase).

  • Effectiveness varies by prompt-image pair: The paper states that although both types of negatives help, their effectiveness varies across prompt-image pairs (p, x), which is what motivates automated selection rather than manual trial and error.

  • Caption grounding helps: The caption-conditioned proposer (NPC) marginally but consistently outperforms a prompt-only proposer on GenEval++ — 0.57 vs. 0.53 accuracy — indicating that grounding negatives in actual image content is more effective than generating them from the prompt alone.

  • Salient score reduces regeneration cost: With the candidate set fixed at five candidates, processing candidates in descending salient-score order lowers the mean number of generation attempts from 4.1 (random order) to 2.5.

  • Salient score correlates with image-level quality: For 204 out of 280 prompt pairs (72.9%), the higher-salience negative produced the image with the higher verifier score.

  • Model-agnostic behavior: Applying NPC improved performance on every evaluated attribute across SDXL, SD3, and FLUX base models. A configuration substituting Qwen2.5-VL (verifier and captioner) and Qwen3 (proposer) achieved performance comparable to the GPT-based setup. In a T2I-CompBench experiment implemented on SDXL, NPC attained the best attribute-binding performance and nearly a 15 percentage-point gain on the complex task, with an average of 0.5237 versus 0.4819 for CoMat.

  • Comparison with a negative-prompt optimization method: On T2I-CompBench, NPC (on SDXL) outperforms DPO-Diff, which scores 0.3746 average versus NPC's 0.5237; unlike DPO-Diff, NPC requires no training.

  • DPG-Bench: NPC using FLUX attains 93.34 relation accuracy and the highest overall score of 85.38, ahead of BAGEL at 85.07, while lagging top systems on Global/Other and slightly trailing on Attribute.

  • Runtime profile: Averaged over 10 randomly sampled cases, the pre-check stage takes about 2.77 s, captioning about 4.55 s, proposal about 0.78 s, and salient scoring about 23.48 s. Scoring is the longest stage; the paper reports that memory changes are only marginal because captioning/proposal use the GPT API and scoring runs on CPU.

Methodology in Plain English

The starting point is a standard text-to-image setup: a positive prompt produces an image, and classifier-free guidance controls how strongly the model follows that prompt. The authors repurpose the "negative" slot of CFG — normally filled with generic quality negatives like "blurry" — to instead hold content that should be suppressed.

The pipeline runs as a generate–evaluate–refine loop. First, a verifier (GPT-4.1 in the main configuration) judges whether the generated image matches the prompt; if it does, the pipeline stops immediately. If not, the verifier returns a short failure reason — for example, wrong attribute — which becomes a targeted negative. Second, a captioner (GPT-4o-mini) writes a natural-language description of the image, enumerating incidental objects and context, which supplies untargeted negatives. Third, a proposer (GPT-4o-mini) combines the failure reason and the caption into five short negative-prompt candidates, typically one or two tokens each.

Rather than regenerate images for every candidate, the authors score candidates in text space. The salient score computes the text-encoder embedding of the positive prompt and of the candidate negative, takes the subtractive direction between them, identifies the prompt's salient tokens via a simple heuristic (the noun following each quantity token), and averages the cosine similarity between that direction and each salient token embedding. Candidates are then tried in descending salient-score order, and the loop stops at the first image the verifier marks correct; if none passes, it falls back to the image with the highest verifier score.

The authors instantiate the image generator with FLUX.1-dev, use 50 denoising steps, apply the negative prompt only during the first three steps (t = 1, 2, 3), set the guidance scale to 3.5 and the true-CFG scale to 1.8, and use an NVIDIA A40 GPU with PyTorch. The paper's attention analysis separates the roles of keys and queries: keys come from the positive prompt and stay fixed, while queries come from the image latent and shift under negative prompting, so changes in the salient-attention score reveal how negatives reallocate the model's semantic focus.

Why This Matters

The work reframes negative prompting as a first-class alignment lever rather than a quality-control afterthought, and it provides an attention-based account of why negatives that are only tangentially related to an error can still help. It also shows that useful negatives can be selected cheaply in text space, avoiding the cost of brute-force regeneration — an important practical point for deployment.

Real-world applications:

  • Content creation and design tooling: automated alignment correction reduces the manual prompt iteration cycle for designers generating marketing or editorial imagery.
  • Creative and surreal illustration: prompts asking for fantastical object transformations (the Imagine-Bench setting) benefit from negatives that preserve identity while enabling the fantasy.
  • Compositional and catalog imagery: prompts specifying counts, colors, positions, and sizes (the GenEval++ setting) map onto product photography and scene mockups where precise attribute control matters.
  • Accessible text-to-image interfaces: because the pipeline is fully automated, users without prompt-engineering skill can obtain better-aligned outputs from natural requests.

Industry relevance: The approach is model-agnostic — gains were observed across SDXL, SD3, and FLUX — and requires no training or gradient updates, only inference-time API calls to an LLM and a text encoder pass. That makes it attractive for teams that want alignment improvements on top of an existing open generator without fine-tuning. The paper also notes that captioning and proposal use the GPT API rather than local GPU compute, which limits added local memory pressure.

Future Directions

  • Closing the remaining category gaps: NPC ranks second best on Color and Color/Pos on GenEval++ and trails T2I-R1 on Spatiotemporal in Imagine-Bench, leaving room to understand which failure modes the negative-prompt approach does not cover.
  • Improving the text-space proxy: the salient score is explicitly stated not to be an exact proxy for image-level effects, and the scoring stage is the longest single stage (about 23.48 s), so a cheaper or more faithful selector is a natural next step.
  • Reducing inference overhead further: since scoring is the dominant cost and negatives are applied only during the first three denoising steps, questions remain about how aggressively the number of candidates or scored steps could be cut without losing alignment gains.
  • Extending beyond the studied negative space: the paper's candidate negatives are typically one or two tokens; whether richer multi-token or compositional negatives yield further gains, and how the verifier/captioner/proposer stack behaves with other model families, are open questions.

Target Audience

This paper is most useful to researchers and engineers working on diffusion- and flow-based text-to-image systems who care about prompt adherence: alignment and controllability researchers, practitioners building generation pipelines that need reliable compositional output, and readers interested in the interpretability of cross-attention as a diagnostic tool for conditioning. It is also relevant to those studying inference-time optimization methods, since NPC achieves its gains without training or gradient updates.

Authors’ abstract

Despite substantial progress in text-to-image generation, achieving precise text-image alignment remains challenging, particularly for prompts with rich compositional structure or imaginative elements. To address this, we introduce Negative Prompting for Image Correction (NPC), an automated pipeline that improves alignment by identifying and applying negative prompts that suppress unintended content. We begin by analyzing cross-attention patterns to explain why both targeted negatives-those directly tied to the prompt's alignment error-and untargeted negatives-tokens unrelated to the prompt but present in the generated image-can enhance alignment. To discover useful negatives, NPC generates candidate prompts using a verifier-captioner-proposer framework and ranks them with a salient text-space score, enabling effective selection without requiring additional image synthesis. On GenEval++ and Imagine-Bench, NPC outperforms strong baselines, achieving 0.571 vs. 0.371 on GenEval++ and the best overall performance on Imagine-Bench. By guiding what not to generate, NPC provides a principled, fully automated route to stronger text-image alignment in diffusion models. Code is released at https://github.com/wiarae/NPC.

Read the original paper