Skip to content
AI.info

Research

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? Overview Research area: Natural Language Processing / Multimodal Language Modeling / Embodied AI evaluat

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
arXiv
2510.16924
Published
2025-10-19
Authors
Zhihui Yang, Yupei Wang, Kaijie Mo, Zhe Zhao, Renfen Hu

AI summary

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?

Overview

Research area: Natural Language Processing / Multimodal Language Modeling / Embodied AI evaluation.

Technical level: Intermediate — the paper is a benchmark-and-evaluation study; understanding the results requires familiarity with language models, vision-language models, and vector embeddings, but the core argument is conceptual rather than mathematically heavy.

Scope: This paper builds and runs a psychological-theory-grounded benchmark to test whether adding vision to language models actually improves their grasp of sensory, physical-world (embodied) knowledge.

What This Paper Is About

Language models learn about the world mostly from text, which can leave them shaky on questions that depend on real sensory experience — things like whether one object is bigger than another, or what happens to a shape when you fold it. Many researchers assume that adding visual input (vision-language models) should fix this gap. This paper tests that assumption directly by comparing 30 state-of-the-art models on two new tasks covering visual, auditory, tactile, gustatory, olfactory and interoceptive knowledge.

Key Contributions

  1. A new embodied knowledge benchmark with two tasks. The authors design SensoryVec (vector-level evaluation of sensory adjectives) and PerceptualQA (multiple-choice perception questions), both structured around the perceptual system framework from psychology, covering visual, auditory, tactile, gustatory, olfactory senses plus interoception.

  2. SensoryVec: 349 sensory word triples and 1,047 sentences. Each triple pairs a word with its synonym and antonym (e.g., small–little–big); the test checks whether a model's vector for the word is closer to its synonym than to its antonym. Words were retained only if they had sensory ratings above 4 in the source datasets, with synonyms and antonyms identified via WordNet and a thesaurus.

  3. PerceptualQA: 1,400 questions across 9 tasks, with 200 questions per visual subtask and 100 questions per other perceptual task. The five visual subtasks are color attributes, colors in nature, geometry and transformations, symbols, and body; non-visual tasks cover auditory, tactile, gustatory and olfactory perception. Claude-3.5-Sonnet was used to generate candidate questions that human annotators then filtered and revised.

  4. A controlled comparison of 30 models across 6 comparable VLM/LM pairs — VisualBERT & BERT, LLaVA-1.6-Vicuna-7B & Vicuna-7B, LLaVA-1.6-Mistral-7B & Mistral-7B, Qwen-VL-Chat & Qwen-7B, Qwen2-VL-7B-Instruct & Qwen2-7B, and Qwen2-VL-72B-Instruct & Qwen2-72B — where each vision model is built directly on its text-only counterpart.

Main Findings

  • No model does well. Performance on SensoryVec ranged roughly from 50% to 70% accuracy, and the best PerceptualQA model, Claude3.5-Sonnet, reached only 69.04%, against a human baseline of 86.00%. Human agreement in the baseline study produced a Cohen's Kappa of 0.69.

  • Text-only BERT won SensoryVec. BERT achieved the best overall accuracy at 72.21% (visual 70.44, non-visual 74.66), ahead of CLIP at 71.06% (visual 75.37, non-visual 65.07). GPT-2 was the weakest at 50.43.

  • Vision-language models show no clear advantage. In PerceptualQA, VLLMs beat their LLM counterparts by an average of only 2.32%: LLaVA1.6-Vicuna-7B 41.64% vs Vicuna-7B 38.25%; LLaVA1.6-Mistral-7B 45.64% vs Mistral-7B 42.96%; Qwen2-VL-7B-Instruct 51.00% vs Qwen2-7B 49.36%; Qwen2-VL-72B-Instruct 63.89% vs Qwen2-72B-Instruct 62.32%. In SensoryVec, VLMs performed comparably or worse than their text-only counterparts.

  • Visual knowledge is the weakest dimension. Models scored far lower on visual than non-visual items — for example Qwen2-VL-72B-Instruct at 54.45 visual vs 87.50 non-visual, and GPT-4o at 59.45 visual vs 91.00 non-visual. The best visual score, Qwen-Max at 61.05%, falls well short of the human visual score of 85.20%, while humans show only a small gap between visual (85.20%) and non-visual (88.00%).

  • Spatial reasoning is the specific failure point. The hardest subtasks were symbols, body, and geometry and transformations. The authors report that errors were not tied to particular shapes, symbol types, body parts, transformation types, or question targets, indicating a systemic deficiency in spatial reasoning rather than missing visual or conceptual knowledge.

  • Vector errors track word form and frequency. Failure examples include antonym pairs that look similar (sugarless–sugary, unwrinkled–wrinkled, undimmed–dim) or are high-frequency (dry–wet, sharp–blunt, open–closed, green–red), which the authors attribute to over-reliance on distributional semantics.

  • Fine-tuning helps but does not close the gap. Supervised fine-tuning raised GPT-4o-Mini from 58.21% to 79.29%, still below the 86.00% human baseline. On the three hardest visual subtasks (V-GT, V-S, V-B), the fine-tuned model averaged 67.26% versus 90.67% for humans. No significant improvement was observed for the Qwen2 series models.

Methodology in Plain English

The authors first built a lexicon of sensory adjectives spanning six sensory categories, then used WordNet and a thesaurus to attach a synonym and an antonym to each word, forming triples. For each triple they wrote three natural sentences — drafted by GPT-4o and cleaned by humans — so that contextualized models could be probed with real sentences while static models like Word2Vec and GloVe could be probed with standalone word vectors. A model passes a triple if its vector for the target word is closer to the synonym than the antonym.

For the question-answering half, they wrote multiple-choice questions that a person could answer through embodied imagination but that are rarely stated outright in text — for example, which direction a number faces after being rotated, or whether elbows sit higher or lower than hips with hands behind the back. Questions were generated with Claude-3.5-Sonnet, filtered and revised by human annotators, and scored by average accuracy over two trials. Seven native-speaking graduate students supplied the human baseline. Finally, they matched each vision-language model to the text-only model it was built on, so the only difference in each comparison was the presence of visual grounding.

Why This Matters

Impact on research: The paper challenges a widely held assumption that multimodal training automatically yields better physical-world understanding. It also argues that existing visual reasoning benchmarks like CLEVR, MMMU and MMBench require image input at inference time, which makes them unsuitable for the direct text-only versus vision-language comparison this benchmark enables. The resource itself is a diagnostic tool for model analysis.

Real-world applications:

  • Embodied agents and robots that must reason about object positions, sizes and transformations to act in physical space.
  • Autonomous driving, where object understanding and spatial inference are safety-critical.
  • Human-computer interaction and multimodal assistants that need to interpret perceptual descriptions rather than only retrieve facts.
  • Haptic and tactile sensing pipelines, which the authors note must fuse touch with vision and language for cross-modal embodied knowledge.

Industry relevance: The work is directly relevant to teams building multimodal assistants and robotics stacks, and it signals that scaling vision-language training alone — including multi-stage alignment pipelines such as Qwen2-VL's — may not deliver embodied competence. The authors tied this to a broader concern: as embodied AI systems are increasingly built around LLMs and VLLMs, their limits in embodied reasoning become a bottleneck for practical deployment. The research was supported by the Tencent Basic Platform Technology Rhino-Bird Focused Research Program.

Future Directions

  • Build richer training data. The authors call for dynamic sensory data such as video sequences, perception-related brain neural data, and feedback from non-human sensors including haptic sensors and motion capture systems.

  • Design new training tasks and architectures that integrate multimodal perceptual information more effectively than current image-text pair training does.

  • Train jointly with embodied agents on tasks involving interaction with simulated or real environments, so models can explore, ground concepts experientially, and acquire causal understanding beyond statistical correlation.

  • Explore the untapped potential of purely textual data. The limitations section notes the authors have not fully explored this avenue and plan to expand the dataset's size and diversity. The paper also leaves open why fine-tuning helped GPT-4o-Mini but not the Qwen2 series.

Target Audience

Researchers and engineers working on multimodal language models, embodied AI, and robotics who need to know where current models actually fail; benchmark and evaluation designers interested in constructing psychology-grounded datasets; and product teams deciding whether adding vision to a language model will meaningfully improve physical-world reasoning. Readers looking for a new model architecture or training method will not find one here — the paper's purpose is diagnosis, not prescription.

Authors’ abstract

Despite significant progress in multimodal language models (LMs), it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models. To address this question, we propose a novel embodied knowledge understanding benchmark based on the perceptual theory from psychology, encompassing visual, auditory, tactile, gustatory, olfactory external senses, and interoception. The benchmark assesses the models' perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. By comparing 30 state-of-the-art LMs, we surprisingly find that vision-language models (VLMs) do not outperform text-only models in either task. Moreover, the models perform significantly worse in the visual dimension compared to other sensory dimensions. Further analysis reveals that the vector representations are easily influenced by word form and frequency, and the models struggle to answer questions involving spatial perception and reasoning. Our findings underscore the need for more effective integration of embodied knowledge in LMs to enhance their understanding of the physical world.

Read the original paper