Research
DEPICT: Scoring Text-to-Image Alignment by Answer Agreement
Overview Research area: Computer vision and vision-language evaluation, specifically metrics that score how well a generated image matches its text prompt (text-to-image alignment). Technical level: A

- arXiv
- 2610.03617
- Published
- 2026-10-02
- Authors
- Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins
AI summary
Overview
Research area: Computer vision and vision-language evaluation, specifically metrics that score how well a generated image matches its text prompt (text-to-image alignment).
Technical level: Advanced. The paper blends probabilistic scoring rules, vision-language model (VLM) prompting, and statistical benchmarking, though its central idea is explained here in plain terms.
Scope: The paper introduces DEPICT, a training-free metric that scores image-caption alignment by measuring agreement between answers a VLM gives from the image and answers it gives from the caption alone, then merges that decomposed score with a holistic whole-caption score, and evaluates it on five benchmarks across eleven backbones.
What This Paper Is About
Text-to-image generators are now good enough that evaluation must catch fine-grained failures: a missing object, a swapped attribute, a wrong count, or an ignored negation such as "no chair." Existing training-free metrics either score the whole caption at once (losing detail) or break the caption into yes/no verification questions (losing context) and grade each answer against a predefined reference, usually assuming the correct answer is always "yes."
DEPICT's goal is to remove that fixed-"yes" assumption. It instead asks whether a VLM's answer based on the image agrees with its answer based on the caption alone, weights each question by how decisively the caption settles it, and combines the result with a holistic score.
Key Contributions
-
Diagnosis of the fixed-yes flaw. The authors identify a structural flaw in current decomposed metrics: fixing the expected answer to "yes" before looking at the image drives their accuracy on negated prompts below chance.
-
Agreement-based scoring. Each question is scored by the agreement between a text-only answer and a visual answer, removing the fixed reference and improving both negation accuracy and correlation with human judgments.
-
Systematic benchmark study. The paper reports what it describes as the first comparison of training-free text-to-image alignment metrics across backbones from several model families and multiple benchmarks, with paired confidence intervals for every difference.
-
DEPICT. A training-free metric that merges agreement-based decomposed scoring with holistic scoring. It exceeds the strongest fine-tuned evaluator on its own backbone on most human-judgment benchmarks and consistently matches or beats training-free baselines across backbones and benchmarks.
Main Findings
-
Negation accuracy rises from 19% to 88%. Replacing the fixed reference with the agreement rule lifts NegBench accuracy from 19% (Soft-TIFA-AM under the Faith rule) to 88.3% (Agree). DSG and Soft-TIFA lose 85–95% of their accuracy on negated items and fall below chance, because 87% of questions for negated captions and 47% for hybrid ones expect "no," against 4% elsewhere.
-
DEPICT beats the strongest fine-tuned evaluator on matched backbone for two of three human-correlation benchmarks. On Qwen3-VL-4B, DEPICT scores 0.630 SRCC on GenAI-Bench versus DynEval-4B's 0.595, and 0.607 versus 0.546 on RichHF. DynEval keeps the lead on TIFA160, 0.802 versus 0.691.
-
A stronger backbone raises training-free metrics substantially. Swapping VQAScore's 2024 backbone for Qwen3-VL-4B, with no other change, raises SRCC by 0.075 on GenAI-Bench and 0.124 on RichHF, enough to surpass every fine-tuned evaluator in the table on those benchmarks, including one trained on a 72B model. Moving DEPICT from Qwen3-VL-4B to Qwen3.5-27B adds a further 0.06–0.10 in every column.
-
Agreement alone closes most of the gap to holistic scoring. On a fixed Qwen3.5-27B backbone, switching from Faith to Agree narrows the gap to VQAScore on GenAI-Bench from 0.15 to 0.03 SRCC, without hurting TIFA160.
-
Merging the two scoring paradigms helps. On all three correlation benchmarks DEPICT matches whichever term suits the benchmark and improves on it. On GenAI-Bench the improvement over the best baseline is +0.022 SRCC, with a paired BCa 95% interval of [+0.016, +0.028]. The two terms disagree on 19% of GenAI-Bench item pairs; DEPICT achieves 60% accuracy on contested pairs, against 53% for the holistic term and 47% for the decomposed term alone. Combining VQAScore with Soft-TIFA instead leaves GenAI-Bench unchanged and lowers NegBench to 62.2, below VQAScore alone.
-
The advantage holds across backbones. Across all eleven backbones, DEPICT matches or beats the strongest baseline in 54 of 55 backbone–benchmark pairs, winning 43 times with 23 of those statistically significant. NegBench gains are larger on InternVL3.5 than Qwen3.5 (4–11 points against 1–4). Its only loss is on GenAI-Bench with the Gemma-4-12B variant that lacks a vision encoder. NegBench accuracy holds at 85–87% at every size on the Qwen3.5 ladder.
-
Pipeline ablations favor direct caption-based questions and no pruning. Tuple extraction costs heavily: 47% of negated NegBench captions yield no tuples and Negation-subset accuracy falls to 4.1, six times below chance. On Winoground pairs, tuple extraction produces overlapping questions (Jaccard 0.44 versus 0.27 without tuples; 13% identical), shrinking the paired score margin by roughly 40% and costing 21 group-accuracy points. Post-hoc pruning lowers performance on every benchmark; commitment weighting handles unreliable questions instead.
-
The scoring rule is the largest single factor. Discretizing Faith as DSG does costs 16.3 Winoground points and 0.033 SRCC, while Faith stays below chance on NegBench. Agree raises NegBench by 4.6 times and GenAI-Bench SRCC by 23%, moving Winoground by at most one point.
-
Cost is roughly doubled in passes but less in wall-clock time. DEPICT requires 2N+2 forward passes per pair (about 12.9 on average), roughly double Soft-TIFA's N+1 (about 6.5). Under continuous batching, measured latency is 0.492 s per pair on GenAI-Bench, 1.36 times Soft-TIFA (0.362 s) and 1.30 times DSG (0.379 s). VQAScore is fastest at 0.099 s with one pass.
-
A stated limitation. On Winoground, whose short captions hinge on compositional word order, decomposition adds no gain beyond the holistic pass, though enabling reasoning on answers improves results there. Because both channels share one backbone, a prior shared by both can make the same hallucination look like agreement.
Methodology in Plain English
The metric runs in four stages, all using a frozen off-the-shelf VLM with no training:
-
Decompose. A language model is prompted with the caption to produce as many yes/no questions as the caption needs, each targeting a single verifiable visual detail. Unlike DSG, which first parses the caption into semantic tuples, DEPICT skips that step, so questions do not inherit parsing errors.
-
Verify. The same VLM answers each question twice: once with the image plus the question, and once with the caption plus the question and no image. Instead of a hard yes/no decode, the method reads the softmax probability of the "yes" and "no" tokens.
-
Agree and weigh. Each question's score is the probability that the two channels give the same answer, summed over both outcomes (both say yes, or both say no). This is symmetric and does not presume "yes." Each question is then weighted by the text channel's commitment, the absolute difference between its yes and no probabilities, so questions the caption settles decisively dominate and ambiguous ones fade.
-
Merge. The weighted average of per-question scores forms the decomposed score. A holistic score is also computed by asking the VLM the whole caption directly (following VQAScore) and reading the probability of "yes." The final DEPICT score is a convex combination of the two, with λ = 0.5 and ε = 10⁻³ fixed in advance and not tuned on any benchmark.
Evaluation setup. Human correlation is measured on GenAI-Bench, TIFA160, and RichHF-18K using SRCC and PLCC, with PLCC computed after a four-parameter logistic fit. Diagnostics use Winoground (text, image, and group accuracy) and the COCO-MCQ split of NegBench. Differences between metrics are paired with BCa 95% intervals over B = 10,000 bootstrap replicates resampled by item. All training-free methods are run on eleven backbones spanning Qwen3.5, InternVL3.5, and Gemma-4, varying family, scale, generation, and architecture, with Qwen3-VL-4B providing a matched comparison to DynEval's backbone.
Why This Matters
Impact on research. The paper reframes how decomposed alignment metrics should grade individual verification questions, showing that the long-standing "yes" target is not a harmless simplification but a structural failure mode that penalizes faithful images whenever the correct answer is "no." It also demonstrates that holistic and decomposed scoring are complementary rather than competing, and provides a cross-backbone comparison protocol with paired confidence intervals that future metric work can reuse. Code will be available for research purposes.
Real-world applications:
- Caption and dataset curation. Scoring image-text pairs to filter noisy training data, as done in large-scale dataset construction.
- VLM hallucination detection. Flagging when a model describes objects or attributes that are absent from the image, including negated content.
- Benchmarking and reward signals for text-to-image generators. Judging whether a generated image renders every aspect of a prompt faithfully, and serving as a reward signal during generation training.
- Negation-sensitive content moderation and accessibility checking. Cases such as "a kitchen with no people" require a metric that correctly rewards faithful absence, which fixed-yes metrics cannot do.
Industry relevance. Because DEPICT is training-free, it can be upgraded simply by swapping in a newer VLM backbone, avoiding the re-training and re-labeling costs that bind fine-tuned evaluators such as ImageReward, PickScore, or DynEval to a fixed architecture and training distribution. The measured latency overhead relative to cheaper baselines is modest (1.36 times Soft-TIFA and 1.30 times DSG under continuous batching), which matters for teams scoring at scale.
Future Directions
- Decoupling the two answer channels. The authors note that using different backbones for the image channel and the caption channel would remove shared priors that can make a mutual hallucination look like agreement, at the cost of hosting two models.
- Closing the gap on short compositional captions. Decomposition brings no gain on Winoground-style prompts that hinge on word order, though enabling reasoning in the answering step improves results there; how to do better remains open.
- Reducing computational overhead. DEPICT needs 2N+2 passes per pair (about 12.9 on average) versus roughly 6.5 for Soft-TIFA; the truncated appendix section on cost begins to address why measured latency overhead is smaller than the pass count suggests, leaving room for further efficiency work.
- Extending the backbone study. The sweep covers eleven backbones from three families; the paper's own observation that InternVL3.5 gains more on NegBench than Qwen3.5 at matched size suggests agreement scoring may compensate for weaker backbones, a question that could be probed with other families and architectures.
Target Audience
Researchers and engineers working on text-to-image evaluation, image-caption alignment metrics, vision-language model hallucination detection, or dataset curation. It is most useful to readers already familiar with baselines such as CLIPScore, VQAScore, DSG, and Soft-TIFA, and with correlation-based benchmarking practice; the plain-language core idea of answer agreement is accessible to a broader technical audience as well.
Authors’ abstract
Image-text alignment is a core problem in computer vision with applications in caption evaluation, hallucination detection, data curation, and the benchmarking of text-to-image (T2I) generators. As T2I models improve, benchmarking has become demanding, requiring metrics capable of finding a series of issues like missing objects, swapped attributes, miscounts, and ignored negations. Recent work addresses this by fine-tuning evaluators on preference data or by prompting a vision-language model, either holistically with the caption or with decomposed verification questions. However, existing approaches fall short: fine-tuned metrics remain bound to one backbone and training distribution; holistic metrics miss fine-grained details; and decomposed metrics rely on a fixed-YES assumption that penalizes faithful images whenever that assumption fails. In contrast, we propose DEPICT, a training-free metric that replaces fixed reference answers with expected agreement between image-based and caption-only answers, weighting questions by how decisively the caption determines them. By replacing fixed references, our agreement rule increases negation accuracy from 19% to 88%. To recover the context lost during decomposition, DEPICT merges this agreement score with a holistic score. We evaluate DEPICT on five benchmarks and eleven backbones from three model families and find that it surpasses all training-free metrics and exceeds fine-tuned evaluators on two out of three human-correlation benchmarks.