Skip to content
AI.info

Research

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning Overview Research area: Computer vision and vision-language research, speci

arXiv
2511.21002
Published
2025-11-26
Authors
Xiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang, Xiaopeng Liu, Min Zhang, Jun Yu

AI summary

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

Overview

Research area: Computer vision and vision-language research, specifically news image captioning, multimodal retrieval-augmented generation (RAG), and knowledge-graph-augmented generation with multimodal large language models (MLLMs).

Technical level: Advanced. The paper assumes familiarity with retrieval-augmented generation, multimodal LLMs such as InstructBLIP, CLIP and InsightFace embeddings, named entity recognition, and graph attention networks.

Scope: The paper introduces MERGE, a multimodal entity-aware RAG framework that pairs an entity-centric multimodal knowledge base with hypothesis-caption alignment and retrieval-driven knowledge-graph construction to generate journalistically informative news image captions.

What This Paper Is About

News image captioning requires describing not just what is visible in a photo but also who and what the photo refers to, which often depends on facts that never appear in the accompanying article. Existing methods fail in three ways: they miss information not stated in the article, they cannot reliably align visual objects with textual specifics such as dates and quantities, and they struggle to attach the right name to the right person or object in crowded scenes. MERGE addresses all three by retrieving external multimodal and structured knowledge at caption time and feeding it, along with a reasoning-derived hypothesis caption and a background knowledge graph, into a fine-tuned multimodal language model.

Key Contributions

  1. MERGE framework: The authors present MERGE, described as the first Multimodal Entity-aware Retrieval-augmented GEneration framework customized for news image captioning, integrating explicit multimodal knowledge with the implicit reasoning of MLLMs.

  2. Entity-centric Multimodal Knowledge Base (EMKB): A knowledge base consolidating named entities, images, and structured background knowledge (sourced from Wikipedia and IMDb), containing 489,085 entities and 2,186,557 images in total. It includes entities extracted from GoodNews and NYTimes800k with spaCy and expanded via an LLM, plus four public datasets (IMDb-WIKI, VGGFace2, CACD, IMDb-Face).

  3. Hypothesis Caption-guided Multimodal Alignment (HCMA): A three-stage Chain-of-Thought prompting mechanism (hypothesis caption generation, relevant sentence selection, global summary generation) for sentence-level cross-modal alignment.

  4. Retrieval-driven Multimodal Knowledge Integration (RMKI): A module with two Retrieval-Augmented Strategies (RAS): entity matching via face embeddings from InsightFace and CLIP image embeddings, and dynamic background knowledge graph construction that combines base relations from spaCy NER and LLM relation extraction with entity subgraphs retrieved from EMKB.

Main Findings

  • State-of-the-art caption quality on GoodNews and NYTimes800k: MERGE achieves CIDEr gains of +6.84 on GoodNews and +1.16 on NYTimes800k over the strongest baseline, EAMA. MERGE reaches CIDEr 94.54 on GoodNews and 88.16 on NYTimes800k.

  • Improved named entity recognition: F1-score improvements of +4.14 (GoodNews) and +2.64 (NYTimes800k) over prior work. On GoodNews, MERGE reaches F1-score 32.40 versus 28.26 for Xu et al. (2024a) and 28.23 for EAMA; on NYTimes800k, MERGE reaches F1-score 33.83 versus 31.19 for Xu et al. (2024a) and 30.97 for EAMA.

  • Precision trade-off on NYTimes800k: MERGE's precision of 31.87 on NYTimes800k slightly trails Xu et al. (2024a) at 32.38; the authors attribute this to that baseline using additional training data (20% training set plus full validation) and knowledge distillation, whereas MERGE's HCMA requires no extra training.

  • Generalization to Visual News: Although Visual News was excluded from EMKB construction, MERGE beats the second-best method, Zhou et al. (2022), by +20.17 in CIDEr and +6.22 in F1-score, reaching CIDEr 127.77 and F1-score 29.66.

  • Zero-shot MLLMs underperform: Without fine-tuning, InstructBLIP scores CIDEr 24.42 and F1-score 15.17 on GoodNews; with fine-tuning these rise to 84.8 and 29.76. Full MERGE reaches 94.54 and 32.40.

  • Ablation evidence for each component: On GoodNews, adding HCMA Stage 1 raises CIDEr from 84.8 to 84.83, Stages 1+2 to 85.24, and Stages 1+2+3 to 86.08; adding RMKI and EMKB with RAS 1 reaches CIDEr 91.52; RAS 1+2 reaches 91.36; the full MERGE reaches 94.54. On Visual News, InstructBLIP (w/ FT) scores CIDEr 103.48, RAS 1 reaches 116.84, RAS 1+2 reaches 123.50, and full MERGE reaches 127.77.

  • Complementary retrieval strategies: RAS 1 (entity matching) primarily improves entity recognition precision, while RAS 2 (background knowledge graphs) improves contextual grounding and recall; combining both yields the highest F1 and CIDEr scores in the ablation.

  • NER gains by entity type: On GoodNews, MERGE reports PERSON precision/recall/F1 of 48.16/42.67/45.25, GPE 31.01/35.61/33.15, and ORG 28.13/28.40/28.26. On NYTimes800k, PERSON 47.09/47.51/47.30, GPE 31.49/43.62/36.58, ORG 26.60/29.24/27.86. On Visual News, PERSON 39.55/39.40/39.47, GPE 25.30/31.79/28.18, ORG 25.89/22.18/23.89.

  • Advantage over off-the-shelf MLLMs: MERGE (CIDEr 94.54 on GoodNews) substantially outperforms GPT-4o (18.88), Claude-3.5-Sonnet (29.05), LLaVA-1.6-7B (15.08), Qwen2.5-VL-7B (20.66), and Qwen2.5-VL-32B (12.78). The authors note GPT-4o and Claude-3.5-Sonnet were evaluated on only 1,000 samples due to budget limitations.

  • Backbone flexibility: The paper reports MERGE with Qwen2.5-VL-7B (CIDEr 89.63, F1 30.94 on GoodNews) and LLaVA-1.5-7B (CIDEr 90.49, F1 31.11) versus InstructBLIP (CIDEr 94.54, F1 32.40); the table's NYTimes800k results beyond the first row are not fully reported in the available content.

  • Qualitative case evidence: MERGE correctly identifies Clint Eastwood using EMKB even though his name is absent from the article, aligns details such as "Senate Commerce and Judiciary committees," "11,232 units," and "80 acres," and distinguishes among multiple individuals in a scene.

Methodology in Plain English

The system works in two retrieval-oriented stages that feed into a caption generator.

First, the authors build a large reference library (EMKB) of entities such as celebrities, locations, landmarks, buildings, organizations, artworks, and products. For each entity they store a Wikipedia image, up to five Google Search images, up to five images from four public face/celebrity datasets, and background knowledge from Wikipedia and IMDb that an LLM converts into structured subgraphs. Unlike a static knowledge graph, these subgraphs are fetched dynamically when a caption is being written.

Second, given an image and its article, the HCMA module runs three prompt-driven steps: it drafts a preliminary "hypothesis" caption of up to 30 words using up to 10 key sentences from the article; it then selects up to five sentences that bridge the hypothesis caption and the image; finally it produces a global summary of the article limited to 100 words to recover broader context that the selected sentences miss.

Third, the RMKI module grounds entities. If faces are detected, InsightFace embeddings are compared by cosine similarity against face vectors in EMKB; if no faces are present, CLIP image embeddings are used to find the closest matching images. Separately, it builds a background knowledge graph: spaCy extracts named entities from the selected sentences, an LLM extracts directed relations among them (kept to three words each, one direction per pair) to form a base graph, and each entity's stored subgraph is retrieved from EMKB and merged in with deduplication.

Finally, the image, hypothesis caption, selected sentences, global summary, matched entities, and knowledge graph are combined and passed to InstructBLIP, augmented with a 4-layer Graph Attention Network to encode the graph. Training minimizes a standard cross-entropy loss over ground-truth captions. Hyperparameters include maximum context length 1,024, maximum output length 50, maximum sequence length 4,096, LoRA rank 16 and scaling factor 16, total batch size 16, learning rates of 3.0×10⁻⁵ (GoodNews) and 2.0×10⁻⁵ (NYTimes800k and Visual News), and random seed 42. The authors also note that MLLMs occasionally produce JSON formatting errors, especially with special characters such as colons, which they handle by re-prompting.

Why This Matters

Impact on research: The work argues that pure visual modeling or article-only context extraction is insufficient for news captioning, and that multimodal RAG with a structured, entity-centric knowledge base is a stronger design. It provides an ablation-backed decomposition showing which retrieval strategy helps caption quality versus entity precision, and it demonstrates cross-dataset generalization to a dataset excluded from knowledge base construction, which addresses a common weakness of retrieval systems.

Real-world applications:

  • Newsroom captioning assistance, where editors need entity-accurate captions for photos of people, events, locations, and products.
  • Media archiving and image search, where captions and entity links improve indexing of large photo libraries.
  • Accessible news delivery, including screen-reader descriptions that convey who and what is in an image, not just visible objects.
  • Content verification and fact-checking workflows, where grounding captions in referenced knowledge can surface entity mismatches.

Industry relevance: The paper is affiliated with People's Daily alongside Hangzhou Dianzi University, Harbin Institute of Technology (Shenzhen), and Peng Cheng Laboratory, and the code is released at https://github.com/youxiaoxing/MERGE. The approach targets a production-relevant problem for publishers, and its use of LoRA fine-tuning and a modular retrieval layer suggests it can be adapted to different MLLM backbones, which the authors demonstrate with Qwen2.5-VL-7B and LLaVA-1.5-7B.

Future Directions

  • Keeping the knowledge base current: EMKB covers 489,085 entities and 2,186,557 images. How the base is updated over time so new people, products, and events are covered is described as an updating mechanism in Appendix B, but the long-term maintenance cost of such a large resource remains an open practical question.

  • Closing the precision gap on long articles: NYTimes800k articles are nearly twice as long as GoodNews, and MERGE's margin over baselines there is smaller; better handling of very long, noisy articles is a natural next step.

  • Standardizing LLM structured output: The authors report JSON formatting errors from MLLMs, particularly with special characters like colons, which require re-prompting. More robust structured generation would remove this failure mode.

  • Extending and stress-testing multimodal RAG: The paper confirms the value of combining visual grounding with structured knowledge retrieval, raising the question of how far this pattern transfers to other knowledge-intensive vision-language tasks beyond news captioning.

Target Audience

This paper benefits researchers and engineers working on multimodal large language models, retrieval-augmented generation, image captioning, and knowledge-graph-augmented generation, as well as practitioners building AI-assisted tools for newsrooms, media archives, and image accessibility. Readers need a background in vision-language models and RAG to follow the technical details, but the high-level framing of the three challenges and the component-ablation results make the paper's central argument accessible to those focused on applied systems.

Authors’ abstract

News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances, existing methods struggle with three key challenges: (1) incomplete information coverage, (2) weak cross-modal alignment, and (3) suboptimal visual-entity grounding. To address these issues, we introduce MERGE, the first Multimodal Entity-aware Retrieval-augmented GEneration framework for news image captioning. MERGE constructs an entity-centric multimodal knowledge base (EMKB) that integrates textual, visual, and structured knowledge, enabling enriched background retrieval. It improves cross-modal alignment through a multistage hypothesis-caption strategy and enhances visual-entity matching via dynamic retrieval guided by image content. Extensive experiments on GoodNews and NYTimes800k show that MERGE significantly outperforms state-of-the-art baselines, with CIDEr gains of +6.84 and +1.16 in caption quality, and F1-score improvements of +4.14 and +2.64 in named entity recognition. Notably, MERGE also generalizes well to the unseen Visual News dataset, achieving +20.17 in CIDEr and +6.22 in F1-score, demonstrating strong robustness and domain adaptability.

Read the original paper