Research
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration Overview Research area: Multimodal large language models (MLLMs), specifically hallucination mitigation in vi

- arXiv
- 2512.02981
- Published
- 2025-12-02
- Authors
- Zhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An, Wei Pang
AI summary
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent CollaborationOverview
- Research area: Multimodal large language models (MLLMs), specifically hallucination mitigation in vision-language tasks (computer vision / NLP).
- Technical level: Intermediate. The paper assumes familiarity with transformer attention, entropy, and multi-agent LLM pipelines, but the framework itself is training-free and conceptually explainable.
- Scope: A single paper proposing InEx, a training-free multi-agent framework that combines internal introspective reasoning with external cross-modal verification to reduce hallucination in MLLMs.
What This Paper Is About
Multimodal LLMs often produce responses that sound plausible but are factually wrong relative to the input image, a problem called hallucination. Existing fixes either require human supervision and retraining, or rely only on the model's internal signals and lack outside verification. InEx instead mimics how humans make reliable decisions: reason internally first to reduce uncertainty, then have the result checked from independent external perspectives before finalizing an answer.
Key Contributions
- A training-free multi-agent framework (InEx) that mitigates hallucination through internal introspective reasoning ("In") and external cross-modal multi-agent collaboration ("Ex"), requiring neither human supervision nor model retraining.
- A demonstration that cross-modal consensus mitigates hallucination, supported by multi-modal, multi-perspective self-reflection agents that verify responses against a different modality than the one that produced them.
- TVER-guided introspection: the paper systematically investigates internal-reasoning-based hallucination detectors and finds Text-to-Visual Entropy Ratio (TVER) to be the most effective, reporting the highest AUROC and lowest ECE (Figure 1), and uses it as the trigger signal for internal refinement.
- Strong empirical results: InEx consistently outperforms prior baselines by 3.8%–16.2% on general-purpose benchmarks and 6.5%–26.7% on hallucination benchmarks (the abstract states 4%–27% gains overall), plus theoretical justification via three theorems on mutual information, conditional entropy, and Information Bottleneck loss.
Main Findings
- POPE improvements: On the MSCOCO subset of POPE with LLaVA-1.5-7B, InEx raises accuracy from 79.83 to 88.73 (+8.9%). It achieves the best average across Random, Popular, and Adversarial settings on all three POPE sub-datasets (MSCOCO: 88.73 accuracy / 86.48 F1; A-OKVQA: 87.32 / 87.34; GQA: 86.72 / 86.81), beating OPERA, ICD, and VCD.
- Other hallucination benchmarks: On HallusionBench with LLaVA-1.5-7B, InEx reaches 44.2 accuracy versus the 41.5 baseline (+2.7), while OPERA, ICD, and VCD all score below baseline. On CHAIR it reduces CHAIR S from 50.0 to 45.1 (−4.9) and CHAIR I from 15.4 to 13.3 (−4.1), while raising Recall from 77.1 to 83.2 (+6.1).
- General-purpose gains: On LLaVA-Bench, InEx is the only method applied to LLaVA-1.5 and improves accuracy from 63.4% to 66.5%. For LLaVA-1.5 it leads on MME-Hall (673.3), MM-Vet (36.00), VizWiz (53.80), and MMBench (67.17), with object-level scores of Existence 199.3 and Count 157.7 and attribute-level scores of Position 137.0 and Color 179.3.
- Cross-model robustness: For Qwen-VL, InEx reaches 677.3 on MME-Hall (+59); for GLM-4V, 732.3 (+35). The paper notes that OPERA does not support Qwen-VL and GLM-4V.
- Ablation shows complementarity: Individually, the "In" introspective module and the textual and visual self-reflection components each improve results (Table 3), but the best performance on all seven evaluated datasets comes from combining all three. The full combination yields POPE 88.73, CHAIR 83.20, VizWiz 53.80, MME 653.8, MMBench 67.17, MM-Vet 36.00, and LLaVA-Bench 66.50.
- Image editing choice matters: Testing four image editing models with different architectures (diffusion-based and FLUX-based), InEx with In-Context Edit gives the best overall performance, and InEx improves on the original model across every editing variant.
- Stable and statistically significant gains: On POPE, InEx has a mean of 88.73% with standard deviation 0.33, versus OPERA 84.14 (0.62), ICD 82.97 (0.41), and VCD 82.60 (0.54). Paired one-sided t-test p-values range from 2.85 × 10⁻³⁷ to 8.72 × 10⁻²⁵, and Wilcoxon signed-rank p = 9.54 × 10⁻⁷, at α = 5%.
- Parameter sensitivity: Perspective count and ensemble size produce stable gains up to a value of four; external collaboration iteration count stabilizes at four. Image guidance peaks at 7.5 and text guidance improves up to 10; editing-agent inference steps improve accuracy up to 100.
Methodology in Plain English
InEx has two halves that feed into each other.
The first half, "In," watches the model's internal attention as it generates text and asks a simple question: is the model paying far more attention to the text than to the image? The paper measures this with the Text-to-Visual Entropy Ratio (TVER), the ratio of entropy over textual attention positions to entropy over visual ones. A high TVER suggests the model is confidently following misleading cues rather than the image. When TVER crosses a threshold, the model does three things: it "revisits" the image tokens and blends them into its feed-forward computation (Self-Introspective Visual Augmentation); it filters out uncertain attention heads at the final layer using a Vision-Enhanced Multi-Head Attention variant; and it compares the original and enhanced prediction logits, then either blends them (when they agree) or uses the enhanced version only to calibrate the original (when they disagree).
The second half, "Ex," is a multi-agent review process. A decision agent produces an initial answer. A textual self-reflection agent checks it against a dense caption of the image; if unsupported, it sends feedback and the decision agent revises. Then an image editing agent modifies the image based on the generated answer, and a visual self-reflection agent checks whether the edited image still matches the original (via a CLIP similarity threshold). If it does not, visual feedback goes back to the decision agent. Verification alternates between text and vision until both modalities agree or a preset iteration limit is reached. Textual reflection itself uses multiple perspectives and ensemble aggregation to mimic a rigorous human self-review.
The paper also offers three theorems arguing that injecting visual evidence increases mutual information between hidden states and visual tokens, reduces conditional entropy of the output, and improves the Information Bottleneck objective, drawing on the Data Processing Inequality and Information Bottleneck literature. No component is trained; all settings are fixed hyperparameters (γ_TVER = 0.55, γ_d = 0.2, γ_CLIP = 0.9, α₁ = 1, α₂ = 1, decoding temperature 0.7, ensemble of three outputs, up to four perspectives), and experiments run on NVIDIA H100 GPUs in FP16.
Why This Matters
Hallucination is the main obstacle to trusting MLLM outputs in settings where being wrong has consequences. This paper's contribution is showing that a substantial chunk of that unreliability can be removed without retraining or human labeling, by combining a cheap internal uncertainty signal with a structured external review loop — and by unifying what prior work treated as separate pre-, in-, and post-processing paradigms into one iterative process.
Real-world applications:
- Assistive technology for blind and low-vision users, where VizWiz-style visual question answering errors directly misinform the user.
- Medical and scientific image reporting, where a model describing findings that are not in the scan is a safety failure.
- E-commerce and catalog QA, where product attribute answers (color, count, position) must match the actual image.
- Robotics and embodied agents, which act on visual scene understanding and cannot afford fabricated object or attribute claims.
Industry relevance: The framework is plug-and-play for existing MLLM deployments — the paper evaluates it on LLaVA-1.5, Qwen-VL, and GLM-4V across 7B to 10B scales — and the statistical testing (low standard deviation, significant p-values) speaks to production reliability rather than cherry-picked gains. The trade-off is inference cost: multi-agent iteration and an image editing step add latency and compute, though the paper does not report those costs.
Future Directions
- Scalability beyond the tested range: The paper evaluates 7B to 10B models; whether the introspection-and-verification loop helps or saturates at larger scales is not reported.
- Cost and latency: No runtime, token, or GPU-hour comparison against the single-pass baselines is reported, even though InEx adds iterative agents and an image-editing step. Quantifying this trade-off is an open question.
- Deterministic stopping and robustness: Performance stabilizes at four iterations and four perspectives, but the paper does not report behavior when the textual and visual agents persistently disagree — that is, what the framework does when cross-modal consensus is never reached before the iteration limit.
- Generalization to other tasks: The evaluation is VQA-centric (POPE, CHAIR, HallusionBench, MME, MMBench, MM-Vet, VizWiz, LLaVA-Bench). Extending cross-modal consensus to video, multi-image, or long-form caption generation is untested here.
Target Audience
Researchers and engineers working on multimodal LLMs who need practical hallucination mitigation without retraining budgets; practitioners deploying vision-language models in high-stakes domains; and graduate students studying uncertainty estimation, multi-agent LLM architectures, or the intersection of cognitive-inspired design and model reliability.
Authors’ abstract
Hallucination remains a critical challenge in large language models (LLMs), hindering the development of reliable multimodal LLMs (MLLMs). Existing solutions often rely on human intervention or underutilize the agent's ability to autonomously mitigate hallucination. To address these limitations, we draw inspiration from how humans make reliable decisions in the real world. They begin with introspective reasoning to reduce uncertainty and form an initial judgment, then rely on external verification from diverse perspectives to reach a final decision. Motivated by this cognitive paradigm, we propose InEx, a training-free, multi-agent framework designed to autonomously mitigate hallucination. InEx introduces internal introspective reasoning, guided by entropy-based uncertainty estimation, to improve the reliability of the decision agent's reasoning process. The agent first generates a response, which is then iteratively verified and refined through external cross-modal multi-agent collaboration with the editing agent and self-reflection agents, further enhancing reliability and mitigating hallucination. Extensive experiments show that InEx consistently outperforms existing methods, achieving 4%-27% gains on general and hallucination benchmarks, and demonstrating strong robustness.