Research
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
Overview Research area: Evaluation methodology for vision-language models used as automated annotators (NLP / trustworthy AI evaluation; the paper appears in the TAE (Trust-AI-Eval) workshop track). T

- arXiv
- 2609.37863
- Published
- 2026-09-29
- Authors
- Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem, Lotem Peled-Cohen
AI summary
Overview
Research area: Evaluation methodology for vision-language models used as automated annotators (NLP / trustworthy AI evaluation; the paper appears in the TAE (Trust-AI-Eval) workshop track).
Technical level: Intermediate. The core idea is intuitive, but the paper leans on inter-annotator agreement statistics (Fleiss κ, bootstrap intervals, Welch t-tests) and on the alternative annotator test, so some background in evaluation methodology helps.
Scope: The paper introduces MIST (the Misleading-Image Stress Test), a 200-item English benchmark of potentially idiomatic phrases shown with an aligned image, a misleading image, or no image, and uses it to test whether thirteen VLM judges are destabilised by context they are explicitly told to ignore.
What This Paper Is About
Protocols such as the alternative annotator test (alt-test) decide whether a language model may replace a human annotator panel, but each returns a single verdict from one prompt, one format and one presentation of the item. This paper asks whether those verdicts survive a change in incidental context — specifically, whether attaching an image to a sentence whose label must be decided from the sentence alone changes what a VLM judge answers. The design makes any label change an error by construction, so the authors can separate using an image from being influenced by what it depicts.
Key Contributions
- MIST, a new stress-test benchmark: 200 English items, each a sentence built around one potentially idiomatic phrase (the target), drawn from 200 distinct compounds. Each item has four labels (Fully Figurative, Weak Figurative, Figurative and Literal, Fully Literal) and is presented with an aligned image (depicting the sentence's sense), a misleading image (depicting the opposite sense), or no image, yielding four conditions of 50 phrases each.
- A label-preserving, multi-modal invariance test run against a judge rather than a task model: the image is never the object of the decision, so the correct answer does not move when the image is swapped; any label change is an error.
- A large-scale judge audit: thirteen VLMs (four proprietary, nine open-weight, spanning 3B to 32B active parameters) evaluated under four prompting techniques (zero-shot, few-shot, chain-of-thought, and few-shot with chain-of-thought), each crossed with an explicit "ignore the image" instruction and a silent arm that never mentions it, producing 20 labels per item and 52,000 labels in total.
- A released dataset of all human annotator and VLM judge labels, hosted at huggingface.co/datasets/naghamo/mist-vlm-judges.
Main Findings
- Presence, not content, moves the judge: Attaching an image changed labels at a rate of 15.9% for an aligned image and 14.7% for a misleading image among the seven judges that pass the alt-test in at least one configuration. Across all thirteen judges the rates were 20.5% aligned and 19.4% misleading. The two rates are close for every judge, never more than five points apart.
- A plain-language instruction moves fewer labels than the image itself: Holding the image fixed and deleting the paragraph that instructs the judge to ignore the image changed 8.8% of labels (mean over the seven judges; 11.6% across all thirteen). Every judge individually exceeded its own prompt-edit rate in both image columns, so telling a judge in words to disregard the image is less effective than not placing it there.
- The changed labels do not follow the image: Across all thirteen judges the aligned and misleading images disagree on 17.1% of (item, prompt, judge) cells, 1,776 of 10,400. Of these, 37% moved toward the sense the misleading image depicts and 63% moved away. Chance is 50%; every alt-test-passing judge fell below it, the highest at 39%, and across all thirteen only one judge exceeded chance, by two points.
- The movement is not a change of reading: Among the seven judges, 71% of the movement stays on the same side of the figurative–literal divide, so the image reshuffles the answer without changing which reading the judge believes.
- Accuracy is unchanged: Agreement with the human majority label was 54.4% with no image, 53.7% with an aligned image and 54.7% with a misleading image, and only 5 of 13 judges lost accuracy under an image. The image changes which items a judge agrees on, not how many.
- Weaker judges move more: The effect was smaller in the seven judges that pass the alt-test than in the six that never do, 15.9% against 25.8%, but present in both groups.
- Which judges pass: Against the informed trio, no judge passed the alt-test in any of the 52 judge-by-prompt cells. Against the blind trio, two judges passed at the slack prescribed for its tier (G3.5-F and Gm3-27B) and seven passed at the more permissive ε = 0.20, while six never passed.
- Human annotators are a hard but unresolved baseline: Six fluent English speakers worked as two disjoint trios. The informed trio reached Fleiss κ = 0.65 and the blind trio 0.57, a difference of +0.08 [-0.05, +0.22] with p = 0.22, so the trios cannot be shown to differ. The aligned-versus-misleading contrast was not statistically resolvable for either trio (blind p = 0.27, informed p = 0.92); at 50 items per condition only differences of roughly 0.15 in per-item agreement or 0.20 in Fleiss κ are resolvable.
- Suppression rather than inattention: Under Few-Shot + CoT, one judge exceeded a 0.5 image-following ratio (Q2.5-7B, 0.60) and six fell below 0.4, consistent with active suppression of the image rather than indifference to it.
Methodology in Plain English
The researchers needed a task where the right answer cannot move even though the surrounding context changes. Idioms supply one: a phrase like "kicked the bucket" can be read figuratively or literally, and which reading is correct depends on the sentence, not on any picture. So they took an existing public dataset derived from the AdMIRe shared task, holding 551 English potentially idiomatic expressions, each with a figurative sentence, a literal sentence, and generated images for both readings. They filtered to compounds for which both readings exist in a compatible form, sampled 200 at random (using random.seed(123)), and kept one sentence per compound, 100 figurative and 100 literal. Because the discarded row's image still exists, each retained sentence can be paired with either reading's image, or with none.
Each item was shown in three inputs: aligned image, misleading image, or no image. Human annotators each saw one condition per expression, so they could not infer the manipulation by comparing versions; the VLM judges saw all three inputs of every item. The annotation guidelines state that the label describes how the written phrase reads in its sentence, that a literal image does not by itself make a phrase Fully Literal, and that anyone distracted by the image should annotate as if it were hidden.
For each judge the authors measured four things: how often a label changed when an image was attached (change rate), in which direction the changes went when the two images disagreed (direction), exact match with the human majority label (agreement), and the alt-test against each trio at its prescribed slack (ε = 0.20 for the informed trio, 0.15 for the blind trio). They also ran a control: keeping the image but deleting only the instruction to ignore it, which bounds how much a prompt edit alone can move a judge.
Why This Matters
Substitutability verdicts are treated as properties of a model, but this work shows they are properties of a model and the configuration it was measured in: the fact that an image is present, not what it shows, shifts labels, and aggregate agreement with humans is blind to the change because the errors cancel out. A practitioner reading "this judge passed the alt-test" would not learn that the judge behaves differently under conditions the guidelines declare irrelevant.
Real-world applications affected:
- Dataset creation and evaluation-set construction: once a human panel is replaced, judge labels become evaluation sets and training data, so instability propagates into flawed comparisons and downstream conclusions.
- Annotation pipelines in multimodal settings: items arrive with images, screenshots, or attachments that guidelines may call irrelevant; a judge may still be destabilised by their mere presence.
- Model and vendor selection: the effect appeared in all thirteen judges across seven families and a tenfold parameter range, so it is not a property of one model or vendor, and the instability was larger in the six judges that never pass the alt-test.
- Documentation of deployment setups: the authors argue a deployment report should state the prompt used and any irrelevant context attached, since the verdict names only the judge.
Industry relevance: for teams that use VLM judges as stand-ins for human annotators, the paper implies that benchmark and validation results should be reported alongside the exact input configuration, and that a judge could pass an invariance test while still being sensitive to trivial context changes.
Future Directions
- Testing truly irrelevant images: both images in MIST depict a reading of the target phrase, so the study contrasts two relevant images rather than a relevant against an irrelevant one. It is untested whether an image about nothing in the sentence would move as many labels again.
- The complementary measurement: whether judges also move appropriately when the sentence genuinely changes reading is left unmeasured by MIST.
- Generalisation beyond this setting: whether the effect is specific to the alt-test and to figurative annotation is posed as the next question.
- Human baselines for change rate: no human labelled the same item twice by design, so the paper cannot say how often a person's label moves for reasons unrelated to the image; the prompt-edit control is only a within-judge substitute. The scale itself is also only partially ordinal, since the FF/WF boundary is a judgement about phrase type and the FL/LL boundary about the instance.
Target Audience
Researchers and practitioners in NLP evaluation who design or rely on LLM/VLM-as-annotator protocols, particularly those working on multimodal judges, benchmark validation, and inter-annotator agreement. It is also relevant to dataset curators who need to know how much weight a substitutability verdict can bear, and to anyone interested in invariance testing as an alternative to aggregate accuracy scores.
Authors’ abstract
Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.