Skip to content
AI.info

Research

"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models

Overview Research area: Human-Computer Interaction, specifically accessibility research at the intersection of blind and low-vision (BLV) technology use and vision-language model (VLM) evaluation. The

"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
arXiv
2511.08917
Published
2025-11-12
Authors
Kapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan, Darren Gergle, Erik B. Sudderth, Anne Marie Piper

AI summary

Overview

Research area: Human-Computer Interaction, specifically accessibility research at the intersection of blind and low-vision (BLV) technology use and vision-language model (VLM) evaluation. The paper carries two ACM CCS categories: "Empirical studies in HCI" and "Empirical studies in accessibility."

Technical level: Intermediate. The paper combines a survey study with quantitative model benchmarking and regression analysis, but the statistical methods are standard (descriptive statistics, Mann-Whitney U tests) and no model training or architecture work is required to follow it.

Scope: A survey of 86 BLV people plus a benchmark of four VLMs on 1,859 naturalistic product images, measuring how blur, misframing, and rotation (singly and in combination) degrade product-identification captions.

The paper was presented at CHI '26 (Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, April 13–17, 2026, Barcelona, Spain), DOI 10.1145/3772318.3791309, arXiv:2511.08917v3 [cs.HC].

What This Paper Is About

BLV people increasingly rely on VLM-based tools such as Be My AI, Microsoft Seeing AI, and ChatGPT to identify everyday products like food, toiletries, and cleaning supplies. These models work well on clean photos, but BLV photographers frequently produce images that are blurry, misframed, rotated, or poorly lit, and no prior work had systematically measured how each of those specific issues changes caption accuracy on real product images. The paper closes that gap by first surveying BLV people about their preferences, photo-taking difficulties, and captioning errors, then building an annotated dataset of their own product photos to test how robust four leading VLMs actually are.

Key Contributions

  1. Empirical evidence of BLV people's preferences and experiences with VLM captioning tools. A survey of 86 BLV participants documents when they choose AI over human assistance, how long photo-taking takes, which image quality issues they perceive as most damaging, and how well current tools help them diagnose those issues.

  2. A disability-centered model evaluation approach. The paper describes the complexities of task and data selection, annotation procedures, and metric/model choices when the evaluation is designed around BLV people's information needs, answering prior calls to address disability bias in AI systems.

  3. A benchmark of four widely used VLMs on degraded real-world product images. GPT-4.1, Gemini 2.5 Flash, Llama 3.2 90B, and Molmo 72B were evaluated on an annotated dataset of 1,859 naturalistic product images derived from the VizWiz dataset, with regression analysis identifying which models are sensitive to which quality issues and content types.

  4. Concrete recommendations across the development pipeline. The paper offers guidance for HCI and ML researchers spanning data curation, model performance improvement, and captioning error handling.

Main Findings

  • High-quality images are handled well, degraded images are not. The best VLM reaches 98% accuracy on images with no quality issues, but accuracy falls to 75% overall once quality issues are present. In the introduction, the authors report accuracy rates of 95% or better for GPT and Gemini on high-quality images taken by BLV people.

  • Compounding issues hurt most. With multiple image-quality issues present, the best-performing model, GPT, drops to 69% accuracy. Figure 1's examples (Campbell's Chunky Chicken Corn Chowder, Kellogg's Corn Pops Cereal, CVS Cortizone cream, Spray 'N Wash Max Laundry Stain Remover, Kraft Deluxe Mac and Cheese, Pillsbury Moist Supreme Devil's Food Cake Mix) were not fully and accurately recognized by any of the four VLMs tested, even though all products are visually recognizable by sighted people.

  • AI is preferred for food, personal care, and unknown household items. Roughly two-thirds of participants said they would almost always or most often use only AI when reading a food label (68.6%, n=59), identifying personal care products or toiletries (67.4%, n=58), and identifying an unknown item in their home (64%, n=55). More than 45% (n=39) would rely on AI to read a medication label despite tool warnings against this use.

  • Humans are preferred for store search and when safety or accuracy matter most. Participants leaned toward human-sighted assistance when searching for a specific product in a physical store (54.7%, n=47) or browsing a physical store (48.8%, n=42), and also when safety (53.5%, n=46) and accuracy (44.2%, n=38) mattered most; cost of searching a large space was a stated reason.

  • Personal privacy pushes people toward AI, data privacy pushes them away. More than half (55.8%, n=48) said they would most often or almost always use only AI when personal privacy matters most, citing embarrassment discussing personal matters with real people. In contrast, 46.5% (n=40) said they would most often rely on human assistance when data privacy was most important, with one participant describing the tradeoff as a catch-22.

  • Taking a good photo remains the hardest step. Two-thirds of participants (67.1%, n=49) said taking a good photo was the hardest part of the captioning process and roughly half (45.9%, n=34) said it took the longest. Most (62.8%, n=54) needed 2–3 photos; 23.3% (n=20) needed just one and 10.5% (n=9) needed more than four. Nearly half (47.7%, n=41) said getting desired information took 2–4 minutes.

  • Framing, blur, and distance are perceived as the most damaging quality issues. On a 4-point scale from "not at all" to "to a great extent," participants rated framing (m=3.54, s=0.71), blur (m=3.5, s=0.69), and distance to object (m=3.45, s=0.6) highest, followed by hand placement/position (m=3.35, s=0.74), lighting (m=3.15, s=0.71), and rotation (m=3.13, s=0.74).

  • Current tools only partially help users diagnose quality problems. Tool support was rated highest for framing (m=2.75, s=0.90), blur (m=2.74, s=0.94), and rotation (m=2.58, s=0.85), and lowest for distance (m=2.22, s=0.82) and hand position (m=2.12, s=0.80), with averages between "Very Little" and "Somewhat."

  • Seeing AI and Be My AI differ in specific ways. Framing was perceived as more impactful on caption quality for Seeing AI (m=3.79) than Be My AI (m=3.39; p=0.0076, U=529.5, n=28 each). Be My AI helped more than Seeing AI with assessing blur (2.92 vs. 2.27; p=0.0147, U=191.5, n=26 and 24) and with hand-obscured products (2.26 vs. 1.81; p=0.0297, U=198.5, n=26 and 23).

  • Feature feedback is mixed. Seeing AI's beep-based framing support drew 14 positive comments but 27 negative or mixed ones, with 21 people saying they had not used it. Be My AI's photo feedback drew 26 positive comments out of 47, with 21 mixed or negative, often citing limited guidance on how to fix a problem.

  • Users are only mildly confident they know why a photo failed. Confidence in understanding why an image was not good enough averaged between "slightly confident" and "somewhat confident" (m=2.39, s=0.97), with only six participants "very confident" or "extremely confident"; no significant difference between Be My AI and Seeing AI.

  • Missing information, not wrong information, is the headline captioning error. The most frequently encountered error in product image captions was missing critical information such as brand names and ingredients, which can be obscured by poor image quality.

  • Tool adoption is broad but human assistance persists. 76.7% (n=66) used AI to identify products at least weekly and 50.0% (n=43) used remote sighted visual interpreting at least weekly. Top tools: Be My AI (76.7%, n=66), Microsoft Seeing AI (69.8%, n=60), AI captioning in screen readers (51.2%, n=44), Access AI (30.2%, n=26), TapTapSee (26.7%, n=23), Google Lookout (11.6%, n=10), WayAround (7.0%, n=6), ChatGPT (38.4%, n=33), Ray-Ban Meta Glasses (29.1%, n=25), Google Gemini (15.1%, n=13), Microsoft Copilot (12.8%, n=11), Claude AI (4.7%, n=4), Clarifai (1.2%, n=1).

Methodology in Plain English

The research proceeds in two stages.

First, the team ran an online survey (hosted on Google Forms, open in March 2025, approximately 10 minutes long, with a $20 Amazon gift card for each participant). Participants were recruited through the research team's email lists plus those of the National Federation of the Blind and the American Foundation for the Blind. Eligibility required identifying as blind or low-vision, being 18 or older, using a screen reader, speaking English, residing in the United States, and regularly using at least one AI tool for image captioning. The survey was iterated three times, including deployments to 10 participants each for feedback. Of 97 responses, 11 were removed for duplicates, quality issues, or invalid email addresses, leaving 86. The survey asked about use of human assistance, accessibility-specific AI tools, and general-purpose AI tools; preferences between AI and humans across scenarios and concerns; and experiences with photo-taking, image quality, and captioning errors. Because the Likert-scale data are ordinal, the authors used Mann-Whitney U tests for comparisons.

Second, the team built an annotated dataset of 1,859 naturalistic product images from BLV people, based on the VizWiz dataset, and evaluated four VLMs — GPT-4.1, Gemini 2.5 Flash, Llama 3.2 90B, and Molmo 72B — on how well they identified the products in those images. They compared accuracy on images without quality issues to accuracy on images with quality issues, examined the effect of compounding multiple issues, and used regression analysis to identify which models are more susceptible to which quality issues and which product content (for example, cans with rounded labels or nutritional facts panels) reduces performance.

The paper frames its evaluation approach as deliberately disability-centered, treating task selection, data selection, annotation procedures, and metric choice as design decisions that shape whose needs the benchmark actually reflects.

Why This Matters

Impact on research. The paper argues that existing caption evaluation methods, which score how well generated text aligns with a reference caption, can produce false positives: a caption can look reasonable while containing serious errors or omitting critical information. By tying model accuracy to specific, real-world image quality issues and to BLV people's stated information needs, the work provides a template for evaluations that center disabled users rather than treating their photos as edge cases — photos that prior work has labeled "other," excluded from analysis, or deferred to future work, despite making up a significant portion of photos taken by blind individuals.

Real-world applications:

  • Accessibility apps and screen readers. Tools such as Be My AI, Microsoft Seeing AI, TapTapSee, and Access AI could use the per-issue findings to target the distortions that most degrade accuracy and to give users actionable, real-time guidance on framing, lighting, and orientation rather than a generic "photo not good enough" message.
  • Grocery shopping and food identification. Because people prefer AI for food labels but humans for in-store search, and because missing brand names and ingredients is the most common error, product-captioning features could be designed to flag when key details are absent.
  • Medication and health-adjacent product reading. Nearly half of surveyed participants would rely on AI to read a medication label despite tool warnings, making robustness on degraded packaging images a safety-relevant concern. Prior work cited in the paper found only 46% of 265 VizWiz medication package images were legible.
  • Photo-capture assistance tools. Systems that coach BLV users through taking photos could prioritize the issues tools currently support least well (lighting, distance, hand position) and move from post-hoc feedback to real-time guidance.

Industry relevance. The four evaluated models are the foundation of many deployed accessibility products, so the finding that the best model drops from 98% on clean images to 75% overall, and to 69% with compounded issues, matters directly to companies shipping VLM-backed assistive features. The paper's recommendations span data curation, model performance, and error handling, giving teams points of intervention across the whole pipeline rather than just the model itself.

Future Directions

  • Diagnosing and fixing model-specific weaknesses. The regression analysis identifies which quality issues individual models are most susceptible to, which the authors frame as a focus for improving performance; the truncated content does not report the regression coefficients, so the precise per-model susceptibilities remain to be read in the full paper.

  • Better error-handling interfaces. Participants reported low confidence in knowing why a photo failed and criticized feedback that said a photo was bad without explaining how to fix it, suggesting a need for real-time and more actionable guidance on framing, lighting, and orientation.

  • Reconciling the privacy tradeoff. Participants split between personal privacy (favoring AI) and data privacy, safety, and accuracy (favoring humans). The paper raises the question of how tools can support one without sacrificing the other.

  • Broader evaluation design questions. The authors describe the complexities of task and data selection, annotation procedures, and choosing models and metrics as open issues for disability-centered evaluation, implying that follow-up work should establish and critique conventions for this kind of benchmark.

Target Audience

This paper is most useful to accessibility and HCI researchers who study assistive technology for BLV people; machine learning researchers and practitioners building or evaluating vision-language models for real-world deployment; designers and product teams shipping captioning or product-identification features; and BLV community members, advocates, and organizations such as the NFB and AFB who want evidence about how these tools perform on the photos people actually take. It is also relevant to researchers developing benchmarks, because much of the paper's argument concerns how evaluation choices encode particular users' needs.

Authors’ abstract

Vision-Language Models (VLMs) are increasingly used by blind and low-vision (BLV) people to identify and understand products in their everyday lives, such as food, personal care items, and household goods. Despite their prevalence, we lack an empirical understanding of how common image quality issues--such as blur, misframing, and rotation--affect the accuracy of VLM-generated captions and whether the resulting captions meet BLV people's information needs. Based on a survey of 86 BLV participants, we develop an annotated dataset of 1,859 product images from BLV people to systematically evaluate how image quality issues affect VLM-generated captions. While the best VLM achieves 98% accuracy on images with no quality issues, accuracy drops to 75% overall when quality issues are present, worsening considerably as issues compound. We discuss the need for model evaluations that center on disabled people's experiences throughout the process and offer concrete recommendations for HCI and ML researchers to make VLMs more reliable for BLV people.

Read the original paper