Skip to content
AI.info

Research

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing Overview Research area: Computer Vision / multi-modal large language models (MLLMs), specifically in

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing
arXiv
2602.04268
Published
2026-02-04
Authors
Siyu Jiang, Feiyang Chen, Xiaojin Zhang, Kun He

AI summary

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing

Overview

Research area: Computer Vision / multi-modal large language models (MLLMs), specifically inference-time methods for reducing visual hallucination during text generation.

Technical level: Intermediate. The paper is readable without prior work on KV-Cache internals, but its core mechanism assumes familiarity with transformer attention, decoding, and hallucination benchmarks.

Scope: The paper diagnoses why MLLM outputs drift away from image content as generation gets longer, and proposes a training-free, plug-and-play smoothing method (KVSmooth) applied to the Key and Value caches during inference.

Authors and venue: Siyu Jiang, Feiyang Chen, Xiaojin Zhang, Kun He, all at Huazhong University of Science and Technology. Posted as arXiv:2602.04268v2 [cs.CV], 11 Mar 2026 (published 2026-02-04), Computer Vision category.

What This Paper Is About

Multi-modal large language models often "hallucinate" — they describe objects, attributes, or relations that are not actually in the image. The authors argue this happens because, unlike text-only models, MLLMs must keep grounding their output in visual evidence, yet the influence of early visual tokens fades as decoding continues, so the generated text gradually drifts away from the image. Their goal is to stop that drift without retraining the model or modifying its architecture.

Key Contributions

  1. A new diagnostic metric called "sink degree." The authors define attention row-entropy as a continuous, real-time way to identify tokens that are accumulating disproportionate attention (so-called sink or aggregation tokens), as an alternative to the slower column-sum metric used in prior work.

  2. KVSmooth, a training-free, plug-and-play inference method. It applies an exponential moving average (EMA) to the Key and Value entries in the KV-Cache, suppressing abrupt state changes during decoding.

  3. Entropy-guided adaptive smoothing strength. Instead of a single fixed smoothing coefficient, the method sets a per-token coefficient from the token's row-entropy percentile rank within a sliding FIFO queue, so high-entropy sink tokens are smoothed more heavily.

  4. Broad empirical validation. Experiments across three MLLMs (LLaVA-1.5, MiniGPT-4, InstructBLIP, all 7B versions) and four benchmarks (CHAIR, OPOPE, AMBER, Object HalBench) show hallucination reductions while preserving or improving coverage of real objects.

Main Findings

  • Hallucinated-object logits grow while ground-truth logits decay. Analyzing 200 images, the authors split objects into GT-InCap (in image and caption), GT-OutCap (in image, not in caption), and Hallucinated (in caption, not in image). Ground-truth objects show a monotonic decline in mean logit with stabilizing variance, while hallucinated objects show rising mean logits with slightly increasing variance, using 95% confidence intervals over twenty generation stages.

  • Row-entropy correlates strongly with attention aggregation. Across all layers, the cosine similarity between attention row-entropy and column-sum has a unimodal distribution centered around 0.79 with low variance.

  • Row-entropy is coupled to hallucination. Cosine similarity between logit ranking and row-entropy is highest for hallucinated objects, meaning more uniform (higher-entropy) attention tracks stronger hallucination tendencies; genuine objects show lower or slightly negative correlations.

  • Smoothing Key and Value together works best. Ablations on CHAIR show joint Key-Value smoothing gives CHAIR_S of 18.2 / 17.0 / 42.2 with F1 of 79.2 / 71.7 / 75.1 (LLaVA-1.5 / MiniGPT-4 / InstructBLIP), versus Key-only (35.6 / 24.0 / 73.2 and 79.4 / 70.6 / 70.2) and smoothing the attention output o_t (33.8 / 19.8 / 84.4 and 74.7 / 66.5 / 61.2). Smoothing raw hidden states caused a severe recall drop.

  • Adaptive coefficients beat the best fixed coefficient. With an optimally chosen constant EMA coefficient ("Ours w/o Ada."), CHAIR_S is 36.2 / 23.0 / 47.8 and F1 is 79.2 / 71.3 / 74.1 — worse than the adaptive version.

  • CHAIR results on LLaVA-1.5: CHAIR_S drops from 41.8 to 18.2 (the paper describes this as a relative reduction of approximately 56%) while F1 rises from 77.5 to 79.2. MiniGPT-4 goes from 31.8 to 17.0 CHAIR_S (F1 69.9 → 71.7), and InstructBLIP from 61.4 to 42.2 CHAIR_S (F1 71.6 → 75.1).

  • Comparison methods are listed for reference. On CHAIR_S for LLaVA-1.5: PAI 22.6, OPERA 44.2, VCD 56.0, SPARC 45.6, MiddleLayer 17.8. For InstructBLIP: MiddleLayer 75.0 versus KVSmooth 42.2.

  • OPOPE: Precision is highest for KVSmooth on all three models (91.20 / 90.74 / 85.19 vs. baselines 86.17 / 87.43 / 83.26), and Fβ=0.2 is also highest (88.90 / 86.59 / 83.69 vs. 85.03 / 83.94 / 82.06). Accuracy is 74.60 / 68.13 / 73.89 against baselines of 76.75 / 67.94 / 73.94. The paper states the method achieves the highest average Accuracy and Fβ=0.2 across random, popular, and adversarial object sets.

  • AMBER and Object HalBench: For LLaVA-1.5, CHAIR (C) falls from 6.1 to 3.1, Hal from 27.3 to 18.7, Cog from 2.8 to 1.3, CHAIR_S from 48.1 to 17.5, CHAIR_SR from 45.3 to 16.7, CHAIR_I from 24.7 to 9.0, while Cover rises from 50.6 to 50.8 and Num from 283 to 286. For MiniGPT-4, C moves from 15.3 to 17.8 and Cover from 63.3 to 59.5, while Hal falls from 65.3 to 39.6 and Cog from 11.0 to 5.1. For InstructBLIP, C falls from 14.7 to 5.4, Hal from 65.8 to 43.7, and Cog from 9.8 to 5.7.

  • Sentence-level hallucination reductions. Object HalBench CHAIR_SR drops by 63.1%, 40.3%, and 41.6% for LLaVA-1.5, MiniGPT-4, and InstructBLIP respectively.

  • Efficiency. The paper reports that KVSmooth keeps inference speed and memory usage close to the baseline while incurring substantially lower overhead than competing methods (detailed in Appendix B, not included in the provided content).

Methodology in Plain English

The method rests on a simple intuition: if the model's internal state changes too violently from one token to the next, it can jump away from what the image actually shows.

  1. Frame smoothing as Bayesian estimation. The authors assume each hidden state equals the previous one plus Gaussian noise. Under this prior, the maximum-a-posteriori estimate of the current state reduces exactly to an exponential moving average — a weighted blend of the new observation and the previous state. The paper notes that σ_o² is unavailable in practice, so the coefficient cannot be derived and must be set another way.

  2. Apply the average to the KV-Cache rather than the hidden state. Key and Value vectors are what produce the hidden state through attention. Smoothing both K and V lets the method use a larger smoothing coefficient and produce the strongest regularization of the output logits, which experiments confirm.

  3. Measure how "sinky" each token is. For each token, the authors compute attention row-entropy across heads and positions — a measure of how diffusely the token spreads its attention. High entropy means the hidden state approximates a historical average and sits near the center of past representations, which attracts more attention later and forms a sink.

  4. Set the smoothing strength adaptively. Recent token entropies are kept in a FIFO queue (length 15 in all experiments). A token's smoothing coefficient is its percentile rank within that queue: the more entropic relative to recent tokens, the more it is smoothed. A reference hyperparameter λ_ref center the values, and the coefficient is clipped to within ±0.2 of λ_ref to avoid extreme suppression.

  5. Plug it in at inference. The approach requires no training, no parameter updates, and no architecture changes. Experiments used greedy decoding with a maximum of 512 generated tokens, applied smoothing to layers 3–31, and set λ_ref to 0.9, 0.5, and 0.7 for LLaVA-1.5, MiniGPT-4, and InstructBLIP respectively — one configuration used across all benchmarks.

Why This Matters

The paper's central claim is that hallucination in MLLMs can be substantially reduced by stabilizing hidden-state dynamics at inference time, without the data and compute costs of fine-tuning or preference optimization. It also reframes sink tokens: rather than only suppressing attention to them, the authors argue these tokens distort internal representations during aggregation and can be regularized directly, which explains the entropy–hallucination link.

Real-world applications:

  • Accessibility tools such as image description or alt-text generation, where describing objects that are not present misleads blind and low-vision users.
  • Medical imaging and diagnostic assistance, where a confidently stated but nonexistent finding is a safety-critical failure.
  • Autonomous driving and robotics, where perception-to-language grounding must reflect what the camera actually captured.
  • E-commerce and content moderation, where automated image captioning and tagging must not invent products, attributes, or policy-relevant content.

Industry relevance: Because KVSmooth is training-free, plug-and-play, and reports inference time and memory close to baseline, it can be applied to already-deployed MLLMs as a decoding-time change rather than requiring a costly retraining cycle — attractive for teams that cannot re-train or re-license foundation models.

Future Directions

  • Automating the reference hyperparameter. λ_ref was set per model (0.9, 0.5, 0.7) and analysis shows larger values give stronger smoothing; a principled way to choose it, or to remove it, is left open.
  • Choosing which layers to smooth. The experiments fixed layers 3–31 across all models; the paper defers investigation of optimal smoothing layers to an appendix, suggesting layer selection is not yet resolved.
  • Approximating σ_o² and σ_p². The derivation shows the ideal EMA coefficient depends on a variance term the authors say is unavailable, so the entropy-based percentile rank is a proxy rather than a Bayesian-optimal solution.
  • Extending beyond object hallucination. Evaluation centers on object presence and captioning benchmarks (CHAIR, OPOPE, AMBER, Object HalBench); whether the same entropy-sink mechanism addresses relational, attribute, or count hallucination is not established in the provided content.

Target Audience

Researchers and engineers working on multi-modal LLMs, hallucination mitigation, or inference-time decoding strategies; practitioners who need to deploy MLLMs in accuracy-critical settings and cannot afford retraining; and graduate students studying attention dynamics, KV-Cache behavior, or the geometry of attention sinks in transformers. Readers need working familiarity with transformer attention and standard vision-language benchmarks to get the most out of the analysis sections.

Authors’ abstract

Despite the significant progress of Multimodal Large Language Models (MLLMs) across diverse tasks, hallucination -- corresponding to the generation of visually inconsistent objects, attributes, or relations -- remains a major obstacle to their reliable deployment. Unlike pure language models, MLLMs must ground their generation process in visual inputs. However, existing models often suffer from semantic drift during decoding, causing outputs to diverge from visual facts as the sequence length increases. To address this issue, we propose KVSmooth, a training-free and plug-and-play method that mitigates hallucination by performing attention-entropy-guided adaptive smoothing on hidden states. Specifically, KVSmooth applies an exponential moving average (EMA) to both keys and values in the KV-Cache, while dynamically quantifying the sink degree of each token through the entropy of its attention distribution to adaptively adjust the smoothing strength. Unlike computationally expensive retraining or contrastive decoding methods, KVSmooth operates efficiently during inference without additional training or model modification. Extensive experiments demonstrate that KVSmooth significantly reduces hallucination ($\mathit{CHAIR}_{S}$ from $41.8 \rightarrow 18.2$) while improving overall performance ($F_1$ score from $77.5 \rightarrow 79.2$), achieving higher precision and recall simultaneously. In contrast, prior methods often improve one at the expense of the other, validating the effectiveness and generality of our approach.

Read the original paper