Skip to content
AI.info

Research

Neural Diversity Regularizes Hallucinations in Language Models

Neural Diversity Regularizes Hallucinations in Language Models Overview Research area: Natural Language Processing (cs.CL) — hallucination mitigation and reliability in language models, combining ense

Neural Diversity Regularizes Hallucinations in Language Models
arXiv
2510.20690
Published
2025-10-23
Authors
Kushal Chakrabarti, Nirmal Balachundhar

AI summary

Neural Diversity Regularizes Hallucinations in Language Models

Overview

Research area: Natural Language Processing (cs.CL) — hallucination mitigation and reliability in language models, combining ensemble theory, portfolio theory, and parameter-efficient fine-tuning.

Technical level: Advanced. The paper's core is a set of formal tail bounds (Lemma 1, Theorem 1, Theorem 2) built on portfolio theory, high-dimensional probability, and econometric terminology. The empirical sections are more accessible, but the contribution is fundamentally theoretical.

Scope: The authors propose "neural diversity" (decorrelated parallel representations) as a provably bounded mechanism for reducing hallucination probability at fixed parameter and data budgets, and demonstrate it concretely with a method called ND-LoRA evaluated on a single small language model.

What This Paper Is About

Language models keep hallucinating even as parameters, compute, and data scale up, and existing fixes (RLHF, RAG, contrastive decoding) mostly optimize average accuracy rather than rare catastrophic failures. The authors reframe hallucination as a second-moment reliability problem governed by representational covariance, prove tail bounds showing that decorrelated parallel streams reduce hallucination probability, and show that excessive parallelism can actually make reliability worse. They then build ND-LoRA — parallel LoRA adapters plus Barlow Twins decorrelation — to demonstrate the theory in practice.

Key Contributions

  1. Theoretical linkage. The first formal tail bounds for hallucination probability in ensembled language models. Theorem 1 shows ℙ(H) ∝ 1/P with P decorrelated parallel representations; Theorem 2 shows non-monotonicity, where excessive parallelism can degrade diversity and therefore reliability. The predictions achieve R² = 0.943, explaining 94.3% of reliability variation across configurations.

  2. Constructive demonstration. ND-LoRA (parallel LoRA + Barlow Twins decorrelation) reduces hallucinations by up to 25.6% and 14.6% on average at 1.00008× continued pretraining cost and 1.1× inference latency, while preserving general capabilities across 12 tasks on Qwen2.5-0.5B.

  3. Mechanistic analysis. Neural diversity is established as a mediator through four routes: causality via perturbation (p < 0.001), quantitative scale via correlation (+0.1% neural correlation associated with +3.8% hallucination), super-linear effects via ablation, and task-dependent optima via scaling sweeps.

  4. Practical approximability. Defaulting to P = 4 achieves 97% of oracle hallucination performance (96% across all 12 evaluations), and a simple prompt-based router achieves 99% of oracle performance.

Main Findings

  • Inverted-U reliability curve. Across P ∈ {1, 2, 4, 8} parallel representations and 6 hallucination benchmarks (182,850 samples, LOWESS, 80% CI), reliability follows an inverted-U, peaking at an optimal P⋆ and then degrading. Hallucination probability correspondingly follows a U-shape.

  • Theory fits empirics with R² = 0.943. The bound uses only two free parameters (C⋆ and SNR) shared across all tasks and observations, with empirical 𝒟(P) plugged in.

  • Task-dependent optima. Hallucination tasks benefit from diversity; knowledge tasks do not. HaluEval-Summ peaks at P⋆ = 4 (0.502, +25.6%), HaluEval-QA at P⋆ = 4 (0.451, +23.4%), HaluEval-Dialog at P⋆ = 4 (0.516, +12.8%), MemoTrap v2 at P⋆ = 8 (0.689, +8.8%), TruthfulQA-MC1 at P⋆ = 2 (0.269, +7.3%), TruthfulQA-MC2 at P⋆ = 2 (0.442, +9.5%), NQ-swap at P⋆ = 8 (0.554, +0.8%). NQ (8-shot, 0.066), PopQA (0.111), and TriviaQA (0.192) all peak at P⋆ = 1 with no reported gains.

  • ≈2× faithfulness advantage. Faithfulness tasks (HaluEval-Dialog, -QA, -Summarization, MemoTrap v2) show roughly twice the benefit of factuality tasks (TruthfulQA-MC1, -MC2).

  • Wins at matched parameters. At P = 2, ND-LoRA R16 beats parameter-matched ParScale R32 (P=2) and Qwen LoRA R32 on HaluEval (0.481* vs 0.439 vs 0.400), MemoTrap (0.666* vs 0.638 vs 0.634), and TruthfulQA (0.442* vs 0.412 vs 0.403), while trailing slightly on NQ (0.055 vs 0.059 vs 0.065) and Wikitext (0.784 vs 0.793 vs 0.778) and slightly winning on Winogrande (0.574 vs 0.564 vs 0.572). HaluEval-Summ improvement is 8.1% absolute, 20.2% relative, p < 0.001.

  • Beats other mitigation methods. ND-LoRA delivers +14.6% hallucination improvement — 3.5× the next-best baseline (CAD, +4.1%) — with knowledge within 0.2% of the P=1 baseline, versus CAD (+1.2%), ActDec (−2.6%), and Disagreement Regularization (−1.1%).

  • Super-linear ablation interaction. Stream LoRA alone (+2.9%) and Barlow Twins alone (+1.4%) sum to 4.3% but achieve +4.9% combined — a 14% bonus. Targeting KVQ attention amplifies this 2.6× to +12.8% at fixed P = 4. Neither component suffices: ParScale's near-complete collapse (𝒟 = 0.9990) yields only +0.5%.

  • Higher 𝒟 can be better. Counterintuitively, ND-LoRA achieves its best result (+12.8%) at 𝒟 = 0.4112, higher than Stream LoRA-BT's 0.1530 — strategic localization to representational bottlenecks matters more than maximizing global decorrelation. Prefix tuning with Barlow Twins collapses streams (𝒟 = 0.9988) while stream-aware LoRA decorrelates them (𝒟 = 0.1530).

  • Causal corruption results. Artificial perturbation of diversity (Δ𝒟 ≈ 0.025) causes significant accuracy drops on HaluEval-Summ (p = 1.6 × 10⁻⁵), MemoTrap v2 (p = 8.2 × 10⁻⁵), and TruthfulQA-MC2 (p = 3.3 × 10⁻⁷), with N = 128 paired samples per task across 4 sub-experiments.

  • Stability across settings. Improvements hold across design layers ℓ⋆ ∈ [7, 23], with λ_BT ∈ [0.01, 0.50] exposing a hallucination–perplexity tradeoff; LoRA rank R16–R128 and alpha scalings are not confounds, and attention-only LoRA outperforms MLP-only LoRA. Even worst-choice P ∈ {2, 4, 8} beats the parameter-matched baseline.

Methodology in Plain English

The authors start by treating a language model's parallel computational streams like a portfolio of correlated assets. Each stream produces a hidden representation that is the "correct" representation plus some noise. If those noises are independent, averaging many streams cancels noise — the same logic behind diversifying a financial portfolio. If the noises are correlated (streams "collapse" into producing near-identical representations), averaging buys nothing. They formalize this with a "neural diversity index" 𝒟: 0 means streams are perfectly orthogonal, 1 means total collapse.

They then derive a bound on the probability that the aggregated output error exceeds a tolerance δ, expressed in terms of 𝒟, the number of streams P, and a signal-to-noise ratio. This gives Theorem 1 (more decorrelated streams → lower hallucination probability) and Theorem 2 (if correlation grows with P — e.g., from optimizer pressure — the hallucination bound becomes U-shaped, so too many streams hurt).

To test this, they build ND-LoRA on top of the ParScale parallel architecture. Each stream gets its own low-rank adapter (rank 16) plus 48 learnable prefix tokens, and the streams are combined with learned weights that include label smoothing (ε = 0.1) so no stream is fully ignored. A Barlow Twins loss is applied across stream pairs at a chosen design layer, pushing their whitened cross-correlation matrix toward the identity — which directly suppresses off-diagonal (cross-stream) correlation. The total loss is cross-entropy plus a weighted Barlow Twins term.

All empirical work uses Qwen2.5-0.5B, continued-pretrained on 20M tokens of The Pile and evaluated on 12 tasks. Comparisons are parameter- and data-matched: ND-LoRA with rank-16 adapters is compared against rank-32 single-adapter baselines. Causal claims come from a corruption hook at the RMSNorm layer that swaps hidden states between streams at random positions, evaluated in a paired design.

Why This Matters

The paper's core claim is that reliability can be improved through representation structure rather than scale, and that tail risk (rare catastrophic failures) is a different objective from average accuracy — the two need not improve together. It provides the first formal tail bounds tying neural diversity to hallucination probability, and the theory's R² = 0.943 fit is unusually tight for hallucination research, where theory typically lags empirics.

Real-world applications:

  • Edge and on-device small language models, which the paper notes are especially vulnerable to hallucination due to compressed representations, and where the 1.1× latency cost matters more than in datacenter settings.
  • Agentic systems where a fabricated fact can trigger an irreversible downstream action, making tail probability rather than mean accuracy the relevant metric.
  • Context-grounded tasks like summarization and document QA, where the paper reports the largest gains (+25.6% on HaluEval-Summarization) because diversity decorrelates internal verification of context-supported claims.
  • Factual retrieval pipelines, where the paper's results suggest diversity does not help — useful negative guidance for practitioners deciding where to deploy the method.

Industry relevance: The method adds +0.008% to continued-pretraining compute and 1.1× to inference latency, using a single frozen backbone with batched adapter kernels — far cheaper than P-model ensembles, which cost P× training and inference.

Future Directions

  • Scaling beyond one small model. All empirical results are on Qwen2.5-0.5B with 20M Pile tokens. Whether the bound and the task-dependent optima transfer to larger models is untested in the material presented.
  • Explaining why higher 𝒟 sometimes wins. ND-LoRA's best result came at 𝒟 = 0.4112 rather than the lowest observed 𝒟 = 0.1530, which the authors attribute to localization at KVQ attention. A principled account of where to place decorrelation, rather than empirical search, remains open.
  • Closing the knowledge gap. Diversity yields no gains on PopQA, TriviaQA, and NQ because it does not add knowledge. How diversity composes with retrieval augmentation or other knowledge-injection methods is not reported.
  • Tuning the diversity–perplexity tradeoff. The λ_BT ∈ [0.01, 0.50] sweep exposes a hallucination–perplexity tradeoff that the paper does not resolve into a selection rule, and the router achieving 99% of oracle performance suggests a learned controller is a viable direction.

Target Audience

Researchers working on hallucination mitigation, LLM reliability, and ensemble or parallel-architecture design; theoretically inclined ML scientists interested in portfolio-theoretic framings of neural computation; and practitioners deploying small language models in edge, agentic, or context-grounded applications where tail risk matters. Readers need comfort with probability bounds and matrix notation to get full value from Section 2, though the experimental sections are readable without it. The paper notes it was reviewed on OpenReview, and code, training and evaluation scripts, and checkpoints are released at github.com/kushalc/nd-lora.

Authors’ abstract

Language models continue to hallucinate despite increases in parameters, compute, and data. We propose neural diversity -- decorrelated parallel representations -- as a principled mechanism that reduces hallucination rates at fixed parameter and data budgets. While existing mitigation strategies largely target accuracy, we provide the first formal tail bounds for hallucination probability in ensembled language models, reframing it as a second-moment reliability problem and explaining 94.3% of empirical reliability variation seen across parallel configurations. We introduce ND-LoRA (Neural Diversity Low-Rank Adaptation), combining parallel LoRA adapters with Barlow Twins regularization, and reduce hallucinations by up to 25.6% (and 14.6% on average) while preserving general accuracy. Ablations show LoRA adapters and regularization act synergistically, causal interventions prove neurodiversity as the mediating factor and correlational studies indicate scale: a 0.1% neural correlation increase is associated with a 3.8% hallucination increase. Finally, task-dependent optimality emerges: different tasks require different optimal amounts of neurodiversity. Together, our results highlight neural diversity as a third axis of scaling -- orthogonal to parameters and data -- to improve the reliability of language models at fixed budgets.

Read the original paper