Skip to content
AI.info

Research

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination Overview Research area: Natural Language Processing / multimodal (omni-modal) large language models, specifically inferenc

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination
arXiv
2610.02999
Published
2026-10-02
Authors
Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue, Guoxin Zhang, Yifan Zhu

AI summary

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

Overview

Research area: Natural Language Processing / multimodal (omni-modal) large language models, specifically inference-time hallucination mitigation.

Technical level: Advanced. The paper uses token-level log-probability contrasts, channel-ablation scoring, and structured attribution, and it assumes familiarity with autoregressive decoding and multimodal benchmarks.

Scope: The paper proposes a training-free, inference-time method that diagnoses which input channel (text, image, audio, video) sustains each generated token, and uses that diagnosis to correct hallucinated outputs, evaluated on a newly constructed 3,540-example benchmark called OmniHalluBench.

What This Paper Is About

Omni-modal large language models (OmniLLMs) unify text, images, audio, and video in one generative model, but they hallucinate when generation relies on the wrong evidence — either leaning on irrelevant evidence or failing to use relevant evidence. Existing inference-time hallucination controls can lower hallucination rates, yet they rarely reveal which specific piece of evidence is sustaining a particular generated commitment. OmniConfess addresses this by freezing a candidate response and re-scoring it under controlled, channel-by-channel evidence removal, producing a token-by-channel "confession" that shows what each token depends on, then using that confession to correct judgments and repair free-form spans.

Key Contributions

  1. A training-free, token- and channel-resolved attribution method. OmniConfess is built on three designs: Commitment Anchoring (generate and freeze a candidate response to eliminate trajectory variation across evidence interventions), Evidence Interrogation (intervene on individual evidence channels and re-score the frozen candidate to isolate token-level, channel-wise dependence), and Confession-Guided Correction (organize these measurements into a token-by-channel confession used to recalibrate judgment scores or locate and revise erroneous free-form spans).

  2. A characterization of evidential dependence in hallucinated OmniLLM outputs. The paper documents three structural properties — front-loaded dependence (the first third of token positions holds 49.8% of dependence mass), cross-channel heterogeneity (coupled channels with heterogeneous directions), and token-level concentration (the top 20% of tokens capture 82% of dependence mass on average) — and uses them to motivate its three designs.

  3. OmniHalluBench, a new benchmark. 3,540 hard examples built from six public datasets, spanning text, image, audio, and video settings and both judgment and free-form generation tasks, including cross-modal conflicts and text-grounded hallucinations.

  4. An empirical study across three backbones and three method families. OmniConfess is evaluated against search-based, evidence-contrast, and fine-grained inference-time baselines on Qwen2.5-Omni-7B (default), Qwen3-Omni-30B-A3B, and Nemotron-3-Nano-Omni-30B, plus ablations of each module and analyses of generation quality, cross-backbone portability, and inference efficiency. Code and benchmark are publicly available.

Main Findings

  • Best overall results on the primary backbone. On Qwen2.5-Omni-7B, OmniConfess outperforms all baselines across all six datasets and all reported metrics. Its average F1 is 65.10 versus 54.80 for the base model, with average Idx 84.94 and average GAV 7.99.

  • Largest gains on PHD and CMM. On PHD and CMM, F1 exceeds the strongest baselines by 10.20 and 9.35 points, respectively (PHD F1 89.60, CMM F1 82.15 on Qwen2.5-Omni-7B). The paper attributes this to channel-wise evidence interrogation curbing reliance on irrelevant channels.

  • Free-form generation improves too. On RAGTruth, Token-F1 gains 16.84 points (55.58), which the paper reads as the confession mitigating underuse of relevant evidence. PubMedQA F1 rises by 2.02 points (97.26) despite a strong baseline.

  • Consistent across architectures. OmniConfess improves average F1 over each base model by 10.30 points on Qwen2.5-Omni-7B (Dense), 6.34 points on Qwen3-Omni-30B-A3B (MoE), and 7.50 points on Nemotron-3-Nano-Omni-30B (Hybrid). Baselines such as AVCD, OPERA, and SC switch between gains and losses across backbones, while GD, PD, and ToT reduce F1 on all three.

  • Every module contributes; correction matters most. Ablating the modules lowers average F1 from 71.63 (full OmniConfess) to 67.06 without anchoring, 63.96 without interrogation, and 58.66 without correction.

  • Anchoring beats a shared prefix. Shared-prefix anchoring raises average F1 from 56.73 to 59.00 and localization AUPRC from 47.28 to 58.20; OmniConfess raises them further to 65.10 and 66.38.

  • Channel-wise confession beats pooled and full-view alternatives. It leads on F1 (82.15) and PhD-Index (84.01), and its Recall Balance score is 86.49 versus 78.16 for pooled evidence, 61.69 for full-view confidence, and 48.84 for the base model — indicating a smaller YES–NO recall gap.

  • Confession-guided revision outperforms confidence-guided revision under an equal correction budget, with mean gains in F1 of +8.92 versus +3.95 and in GAVIE-Acc of +1.49 versus +0.58.

  • Highest judged generation quality. Under cross-evaluation by GPT-4o and human experts across seven dimensions, OmniConfess achieves the highest Overall score (88.7), with Evidence Grounding (86.7), Factual Correctness (92.0), and Uncertainty Calibration (82.1) exceeding the strongest baselines by 5.4, 6.9, and 18.3 points.

  • Latency is the main cost. OmniConfess takes 5.43 seconds on PHD and 35.14 seconds on CMM, versus 2.23 and 20.64 seconds for the base model. CMM generates only 24.6 tokens, suggesting repeated channel-wise rescoring — not output length — drives much of the latency. Compared with search-based methods, OmniConfess uses about 27× and 22× fewer generated tokens on PHD and CMM, but cuts time by only 6.4× and 1.1×. On AVHBench, generated tokens fall from 569.5 to 164.1 and time from 36.15 to 7.93 seconds.

  • Bigger backbones do not guarantee better grounding. In a case study on PHD, OPERA succeeds on the Dense backbone, but its longer reasoning on the 30B backbones narrows the target category and overrides visual evidence.

Methodology in Plain English

The pipeline has three stages.

First, commitment anchoring. The model generates a candidate answer normally and that answer is frozen — both its tokens and its prefix contexts. This matters because the paper shows early tokens carry much of the evidence dependence; if you changed the evidence and let the model regenerate, an early token change would ripple through every later decision, mixing up "the evidence changed" with "the text changed."

Second, evidence interrogation. The system takes the frozen candidate and re-scores it while removing one evidence channel at a time (for example, removing the audio but keeping the image and text). Because only the evidence changes and the text stays fixed, the score differences for each token can be attributed to that channel. This yields per-token, per-channel dependence values and margin contributions, which are organized into a structured "confession" table with diagnostic labels (aligned, under-reliant, over-reliant, dominant). Task rules define which channels are relevant for a given question and task type, so the confession can separate residual preference, relevant evidence, and irrelevant evidence.

Third, confession-guided correction. For judgment tasks, the confession produces a single adjustment applied to the candidate class score, and the model re-decides between the candidate and its fixed alternative (retaining the candidate on a tie). For free-form generation, three confession signals are combined into a per-token revision risk; tokens above a threshold are expanded to word boundaries, merged into contiguous spans, and only those spans are regenerated using the original evidence plus the confession. Content outside the selected spans is preserved.

Prompts are also task-aware: they enforce reliance on provided evidence, require audible-only judgments in CMM, "Yes/No/Maybe" answers in PubMedQA, acknowledgment of absent visual details in HaloQuest, and separate audio-visual observations in AVHBench. No additional training is required, and all experiments run on 4 NVIDIA A40 GPUs.

Why This Matters

Impact on research. The paper shifts hallucination mitigation from a global "is this answer wrong?" question to a local "which evidence, at which token, is doing the work?" question. Its three structural findings (front-loaded dependence, cross-channel heterogeneity, token-level concentration) are reusable diagnostics for anyone studying evidential attribution in multimodal models, and OmniHalluBench gives the community a shared 3,540-example testbed spanning four modality settings and two task types. The negative results about existing paradigms are also useful: search-based methods are task-sensitive, evidence-contrast methods stay close to the base model, and fine-grained token-risk signals apparently cannot tell whether evidence dependence should be strengthened or reduced.

Real-world applications (the paper does not deploy these; they follow from the tasks and modalities it evaluates):

  • Clinical and biomedical question answering, given the PubMedQA text setting where judgments must be evidence-grounded.
  • Retrieval-augmented generation, where RAGTruth-style outputs need span-level auditing of unsupported text.
  • Audio-visual assistants, where an answer may be sourced from sound, sight, or both, and users need to know which.
  • Accessibility and media description, where describing visual details that are absent is exactly the failure mode HaloQuest targets.

Industry relevance. The method is training-free, so it can be layered onto an existing checkpoint without a fine-tuning pipeline. The structured confession is also a plausible audit artifact: it exposes per-token evidence dependence that could feed user-facing citations, compliance review, or human-in-the-loop correction. The efficiency numbers are the trade-off to weigh — latency roughly 2.4× the base model on PHD and 1.7× on CMM, with very low token throughput on audio-visual CMM (0.7 tokens/sec).

Future Directions

  • Reducing interrogation latency. Repeated channel-wise rescoring, not output length, appears to drive cost; the paper notes substantial latency remains on some tasks. Caching, selective channel probing, or batching are unexplored here.
  • Automating channel-relevance rules. The confession currently depends on task rules that define which channels are relevant per question and task type. Whether this relevance signal can be inferred rather than specified is not reported.
  • Scaling the benchmark and modality coverage. OmniHalluBench contains 3,540 examples from six datasets across three backbones; broader coverage and harder conflict cases would test whether gains hold.
  • Closing the gap on the strongest remaining weak spot. AVHBench F1 (22.87 on Qwen2.5-Omni-7B) is far lower than the judgment-task numbers, and the paper reports no comparison against human ceiling performance or against larger closed models.

Target Audience

Researchers and engineers working on multimodal or omni-modal LLMs and hallucination mitigation; practitioners who need training-free, auditable correction methods that can be attached to an existing model; and benchmark designers interested in evidence attribution and hallucination localization. Readers should be comfortable with autoregressive decoding, log-probability scoring, and modality-ablation evaluation.

Authors’ abstract

Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at https://github.com/RongHuiQiang/OmniConfess.

Read the original paper