Research
Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
Overview Research area: Mechanistic interpretability and hallucination mitigation for audio-visual large language models (AVLLMs), with a focus on inference-time, training-free intervention methods. T

- arXiv
- 2609.37568
- Published
- 2026-09-29
- Authors
- Yu Zhang, Pingrui Zhang, Xuefeng Bai, Pengfei Zhang, Yang Xiang, Kehai Chen
AI summary
Overview
Research area: Mechanistic interpretability and hallucination mitigation for audio-visual large language models (AVLLMs), with a focus on inference-time, training-free intervention methods.
Technical level: Intermediate. The paper combines conceptual accessibility (a clear failure mode and a training-free fix) with mechanistic techniques that require some familiarity with transformer internals, such as attention-path cutting, logit lens, and activation steering.
Scope: The paper diagnoses a question-relay mechanism behind source-confused grounding hallucination in AVLLMs and proposes SECRET, a training-free steering method that corrects the failure at question-token positions, evaluated on two benchmarks across three models.
What This Paper Is About
Audio-visual large language models can be misled by the wrong sense: a visible piano in a video makes the model claim it hears piano music even when the audio contains only human speech. This is called source-confused grounding hallucination, where cues from a modality the question did not ask about override evidence from the modality it did. The paper asks where inside the model this interference enters and whether it can be corrected without any retraining.
Key Contributions
- Identification of a question-relay mechanism. Through path-intervention and representation analyses, the authors show that question-token states act as an intermediate relay that carries both required-source evidence and interfering cues from the non-required modality into answer prediction.
- A training-free mitigation method, SECRET (Source-Conditioned Relay Steering). SECRET builds source-conditioned positive and negative question representations by cutting interfering-modality and required-modality attention pathways into question states, then steers the original question states using their norm-matched, token-wise difference.
- Demonstrated effectiveness across models and benchmarks. SECRET improves overall accuracy over base models by up to +18.0 and +7.1 percentage points on CMM and AVHBench respectively, and outperforms the evaluated training-free methods on three AVLLMs (VideoLLaMA2-AV-7B, Qwen2.5-Omni-7B, Qwen3-Omni-30B-A3B).
- Generalization to open-ended generation. Under mismatched audio-video inputs, modality-specific captioning shows lower distractor-reference overlap (D-CIDEr) and higher LLM-score for SECRET across both tested models and both captioning tasks.
Main Findings
-
The model knows which modality is required. Qwen2.5-Omni-7B correctly identified the required evidence modality (audio, video, or ambiguous) for 99.85% of questions when given question text alone. Source-confused hallucination is therefore not a failure to understand what the question asks for.
-
Question states relay required-source evidence. On source-faithful cases in both AVHBench subsets, cutting the attention pathway from required-modality tokens to question tokens (
X_r → X_Q) produced the largest decrease in target-answer probability among the three cuts tested, exceeding cuts to the generation position (X_r → X_G) and to the interfering modality (X_r → X_r̄). This is the basis for the term question relay. -
Interfering cues do enter that relay. Using logit lens over Layers 10–20, cutting the pathway from the interfering modality into question states (
X_r̄ → X_Q) lowered the target-object score in over 90% of examples at every tested layer, approaching 100% at several layers. Mean target-object scores were also lower after the cut. -
Correcting at the question beats correcting at the generation position. Comparing Question Cut (
X_r̄ → X_Q) against Generation Cut (X_r̄ → X_G) over Layers 10–20, the question cut yielded markedly larger correct-answer logit recovery at every tested layer. This contrasts with prior work that targets generation positions. -
Strong benchmark gains. On CMM, SECRET reached 89.3 (+13.4) for VideoLLaMA2-AV-7B, 86.4 (+18.0) for Qwen2.5-Omni-7B, and 87.7 (+8.5) for Qwen3-Omni-30B-A3B overall accuracy. On AVHBench, overall accuracy was 81.0 (+3.6), 84.0 (+7.1), and 81.4 (+4.6) respectively. SECRET achieved the highest overall accuracy among evaluated methods on both benchmarks for each backbone.
-
Gains vary with the required modality. VideoLLaMA2-AV-7B and Qwen2.5-Omni-7B improved more on audio-required tasks, while Qwen3-Omni-30B-A3B benefited more on video-required tasks; the authors attribute this to differences in baseline capability and susceptibility to cross-modal interference.
-
The intervention design matters more than the steering frame alone. In the RQ1 comparison on CMM, SECRET outperformed Gen-Interv. (contrasts built at
X_Gand steered there) and Removal-Interv. (contrasts from modality-removed inputs, steered at question positions) on both VideoLLaMA2-AV-7B and Qwen2.5-Omni-7B. Removing norm matching (w/o Norm) also reduced accuracy on both models. -
Optimal steering depth is model-specific and tracks representation separation. Among tested depths (Layers 21–25), VideoLLaMA2-AV-7B performed best at L₁ = 21 (89.3%) and Qwen2.5-Omni-7B peaked at L₁ = 25 (86.4%). For both models the best-performing depth coincided with the lowest cosine similarity between positive and negative question representations — which increased with depth in VideoLLaMA2-AV-7B and decreased in Qwen2.5-Omni-7B.
-
Weaker cross-modal leakage in captioning. On audio-target captioning, SECRET lowered D-CIDEr from 5.17 to 4.59 and raised LLM-score from 3.12 to 3.76 for Qwen2.5-Omni-7B, and lowered D-CIDEr from 20.2 to 12.8 and raised LLM-score from 2.93 to 3.57 for VideoLLaMA2-AV-7B.
-
Longer clips still improve. In duration-stratified CMM evaluation, gains remained substantial on longer clips, reaching 14.7 percentage points for VideoLLaMA2-AV-7B in the longest reported group (14–16 seconds).
Methodology in Plain English
The authors begin by ruling out the obvious explanation: they check whether the model simply misunderstands which modality the question targets, and find it almost always does understand (99.85%). So they look inside the model instead.
They use attention-path cutting, a technique that severs the attention connections between one group of tokens and another over a window of layers, then observe how the answer's probability changes. Cutting the path from the required modality into the question tokens hurts correct answers the most, showing that question positions carry evidence forward. Cutting the path from the interfering modality into those same question tokens reduces the wrong-object signal that had leaked in, and it restores the correct answer's logit more effectively than cutting the corresponding path into the final generation position.
Based on this diagnosis, SECRET runs three parallel copies of the model through the first L₁ layers with shared parameters. The negative branch blocks the required modality from reaching question tokens; the positive branch blocks the interfering modality instead; the original branch runs unmodified. The per-token difference between the positive and negative question states is rescaled to match the original state's L2 norm and added to it. Those corrected question states are spliced back with the original audio and video states, and the remaining L₂ layers generate the answer. Steering happens only during prefill, with keys and values cached for autoregressive decoding. No parameters are updated at any point.
Why This Matters
The paper shifts the mitigation target for cross-modal hallucination from the generation position, where prior work focused, to intermediate question-token states. It also shows that mechanistic analysis can directly produce a practical, training-free intervention that preserves the full audio-visual input rather than perturbing or removing it.
Real-world applications cited or implied by the paper:
- Autonomous driving, listed as a real-world AVLLM application where grounding in the correct modality is safety-relevant.
- Human–computer interaction, also listed in the paper, where spoken and visual context are mixed.
- Media and accessibility tooling, such as audio description or captioning of video, where the model must describe one modality while ignoring a mismatched other one — the modality-specific captioning setting the paper tests.
- Assistive and monitoring systems that interpret everyday scenes involving people, animals, machinery, and nature, the domains covered by AVHBench's source datasets (VALOR and AudioCaps).
Industry relevance: Because SECRET is training-free and operates at inference time with only attention masking and a state edit, it is a plausible add-on for deployed AVLLMs without retraining or preference-data collection. The rough tripling of the D-CIDEr reduction in captioning (VideoLLaMA2-AV-7B: 20.2 to 12.8) is directly relevant to any product where confidently describing the wrong sense is worse than describing nothing.
Future Directions
- Disentangling modality-specific cues inside question representations, which the authors propose as a way to further reduce source-confused grounding hallucination and use each modality's information more effectively.
- Finer-grained mechanistic analysis, such as examining the roles of individual attention heads rather than only pathways, to clarify the underlying mechanism.
- Generalizing the relay finding beyond audio-visual settings, since the paper's analysis is limited to cross-modal information flow at the pathway level in AVLLMs.
- Working out why optimal steering depth and modality-specific gain patterns differ across backbones, given that the best depth coincided with the lowest positive–negative representation similarity but landed at different layers for VideoLLaMA2-AV-7B (L₁ = 21) and Qwen2.5-Omni-7B (L₁ = 25).
Target Audience
Researchers and engineers working on multimodal or omni-modal large language models, hallucination mitigation, and mechanistic interpretability. It is most useful to readers who want a training-free inference-time method that is grounded in an internal analysis rather than a purely empirical recipe, and who are comfortable with attention-path intervention, logit lens, and activation steering. Practitioners deploying AVLLMs in settings where the wrong modality could mislead a response will also find the evaluation setup and results directly relevant.
Authors’ abstract
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confused grounding hallucination}$, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a $\textbf{question-relay}$ mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose $\textbf{SECRET}$ ($\textbf{S}$ourc$\textbf{E}$-$\textbf{C}$onditioned $\textbf{RE}$lay s$\textbf{T}$eering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.