Skip to content
AI.info

Research

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations Overview Research area: Natural Language Processing (NLP) — LLM hallucination evaluation and computation

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations
arXiv
2512.21635
Published
2025-12-25
Authors
Chengxu Yang, Jingling Yuan, Siqi Cai, Jiawei Jiang, Chuang Hu

AI summary

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations

Overview

Research area: Natural Language Processing (NLP) — LLM hallucination evaluation and computational creativity assessment.

Technical level: Intermediate. The paper assumes familiarity with LLM prompting (SCP, RCP, CoT, RAG), retrieval-augmented generation, and LLM-as-a-judge evaluation, but its core argument is conceptual and accessible.

Scope: The paper introduces HIC-Bench, a benchmark that separates LLM hallucinations into "intelligent" and "defective" types and measures how they trade off across ten scientific domains, six models, and four prompting strategies. arXiv:2512.21635v2 [cs.CL], 30 Dec 2025, by Chengxu Yang, Jingling Yuan, Siqi Cai, Jiawei Jiang, and Chuang Hu.

What This Paper Is About

Most work treats LLM hallucinations purely as errors to be eliminated, yet some "hallucinated" content — such as protein structures later validated experimentally, or the "Learning from Hallucinations" paradigm for robot path planning cited in the paper — turns out to have creative or scientific value. The problem is that nobody has quantified this distinction: existing detectors and benchmarks such as TruthfulQA, UHGEval, HalluDial, and HaluEval measure factual consistency only, so they cannot separate harmless error from productive speculation. The paper's goal is to build a benchmark that classifies hallucinations as either Intelligent Hallucinations (IH) or Defective Hallucinations (DH), measures both, and tests whether reduction strategies can suppress the bad kind without killing the good kind.

Key Contributions

  1. A reframing of hallucinations as a resource rather than only a defect. The authors propose extracting the "most effective hallucination component" in LLMs and treating it as a valuable input for scientific innovation, departing from the conventional "reduce hallucinations" paradigm.

  2. HIC-Bench (Hallucination & Innovation Classification Benchmark). A comprehensive benchmark combining scientific creativity assessment with IH/DH hallucination analysis, built on a dataset spanning ten scientific domains. Three stated core characteristics are structured IH/DH assessment, cross-domain applicability, and dynamic prompt optimization. Code and dataset are released at https://github.com/chujiguangniao/HIC-bench.

  3. The Dynamic Hallucination Prompt (DHP) pipeline. A method that iteratively refines positive and negative example sets based on real-time evaluation feedback, aiming to mitigate the defective aspect of hallucinations while enhancing the intelligent aspect.

  4. An empirical finding that IH and DH are not simply correlated. Evaluation results indicate the relationship is nonlinear and modulated by multiple factors, suggesting that excessive creativity need not be the price of factual reliability.

Main Findings

  • IH and DH are not a simple trade-off. The relationship between intelligent and defective hallucinations "does not exhibit a simple positive correlation but appears nonlinear" — the central claim that creativity and correctness can be jointly optimized.

  • Lower temperature cuts both DH and IH. Comparing generation temperature 1.0 against 0.4 under the SCP prompt, DH fell in all models — gpt-4o-mini declined from 5.00% to 4.50%, which the paper calls the most significant change — but IH also fell, with deepseek-v3 dropping from 13.40% to 12.40%. IFS scores stayed largely stable, so temperature mainly redistributes IH and DH rather than changing the overall balance. deepseek-r1 was excluded from this comparison because it does not support temperature settings.

  • Creativity varies by discipline. Under SCP prompts, most models showed higher Originality in Biomedical Sciences and Aerospace and lower scores in Environmental Science, Social Sciences, and Energy Technology, with IH proportions following a broadly similar trend. deepseek-r1 consistently outperformed the other models on both Originality and IH.

  • Model performance spans a wide range. Under SCP, deepseek-r1 reached 52.60% IH with 1.20% DH and an IFS of 50.04%, while gpt-4o-mini reached only 2.90% IH with 5.00% DH and an IFS of 38.58%.

  • Chain-of-Thought reduces DH and raises IFS but can suppress IH. For gpt-4o, CoT moved IH from 9.60% to 9.30% and DH from 1.30% to 0.40%, lifting IFS from 41.40% to 41.70%. The paper notes the model response varies by model.

  • RAG trades creativity for reliability. Grounding responses in the CDKB knowledge base reduced DH effectively but also constrained IH, which the authors describe as a trade-off between reliability and creativity in open-ended scientific tasks.

  • Relaxing constraints improved both metrics. RCP boosted IH across models while unexpectedly also reducing DH. For deepseek-r1, RCP produced 71.10% IH with 0.30% DH and the highest IFS in Table 2 at 54.10%; for deepseek-v3, 22.20% IH with 2.10% DH and IFS 43.60%; for gpt-4o, 18.60% IH with 0.50% DH and IFS 43.52%.

  • DHP delivered the best combined results. With DHP positive prompting alone, DH dropped to 0.90% for gpt-4o-mini and 0.10% for gpt-4o while IH increased moderately. Combined with RCP constraints, IH rose to 11.10% and 21.10% respectively, with DH at 1.70% and 0.30%, yielding the highest IFS scores in the ablation: 41.54% for gpt-4o-mini and 44.10% for gpt-4o, both above their SCP baselines.

  • The IFS weighting changes which model looks best. Under IIFS (w₁ = 0.9, creativity-focused) and BIFS (w₁ = 0.6, balanced), deepseek-r1 scores highest, but under RIFS (w₁ = 0.1, accuracy-focused) it becomes less suitable due to difficulty maintaining precision.

  • Human verification is built into the loop. Multiple LLM judges score outputs and their scores are averaged to mitigate bias, while human annotators verify the IH versus DH classifications. Automated evaluation used gpt-4o and deepseek-v3 as separate judges to reduce self-preference bias.

Methodology in Plain English

The researchers built a dataset of 100 open-ended innovation tasks spanning ten fields: Quantum Physics, Artificial Intelligence, Biomedical Sciences, Environmental Science, Materials Science, Energy Technology, Neuroscience, Information and Communication Technology (ICT), Aerospace, and Social Sciences. Questions were designed with "fuzzy factual boundaries" so that models would produce hypothetical, predictive, or analogical content rather than straightforward recall. Task construction blended principles extracted from Wikipedia with real-world frontier challenges, and responses were generated ten per task, giving 6,000 responses across six models and ten domains.

Six models were compared to cover different design choices: gpt-4o-2024-11-20, gpt-4o-mini, qwen2.5-14b-instruct, qwen2.5-72b-instruct, deepseek-v3, and deepseek-r1.

Each response was scored on Originality, Feasibility, and Value using a 5-point Likert scale, following the Torrance Tests of Creative Thinking (TTCT) framework. Fluency was measured by SimCSE similarity across outputs and Flexibility by variance in scores across questions. A response counts as an Intelligent Hallucination if Originality ≥ 4, Value ≥ 4, and Feasibility ≥ 3 — high innovation and worth despite partial factual divergence, with plausible realizability retained. Defective Hallucinations are outputs with factual errors, logical inconsistencies, or severe violations of scientific principles; their consistency was checked with kwPrec, a keyword-segmentation metric. Because IH_ratio + DH_ratio ≠ 1, neutral outputs that are neither innovative nor erroneous fall outside both categories.

The two ratios combine into the Intelligent-Fidelity Score, IFS = w₁ × IH_ratio + w₂ × (1 − DH_ratio − IH_ratio), where w₁ and w₂ sum to 1. The main evaluation used w₁ = 0.6 and w₂ = 0.4, prioritizing intelligent hallucinations while ensuring fidelity. Generation temperature was set to 1.0 for creativity with scientific rigor, evaluation temperature to 0, and maximum token length to 70, with Strict Constraint Prompt (SCP) and Relaxed Constraint Prompt (RCP) as the two foundational strategies, plus CoT and RAG layered on the SCP baseline. (The appendix prompt template labels RCP as "Role-Constrained Prompt," while the main text calls it "Relaxed Constraint Prompt.")

DHP works as an iterative loop: human experts seed the process with positive examples that scored high on TTCT metrics and negative examples with high DH rates. During machine evaluation, any response scoring above the initial positive examples becomes a new positive example, and negative examples are continuously refreshed with the latest DH responses.

Why This Matters

The paper reframes hallucination research from pure error suppression to selective cultivation. If IH is a genuine, measurable signal of creative reasoning rather than noise, then hallucination mitigation pipelines that flatten all deviations may be discarding scientific value alongside the errors. The finding that IH and DH are nonlinearly related gives researchers a reason to look for operating points that improve both, rather than accepting an accuracy-versus-creativity trade-off as inevitable.

Real-world applications suggested by the paper's framing and cited examples:

  • Scientific discovery and hypothesis generation, where speculative but plausible concepts can seed new research directions rather than being filtered out as errors.
  • Protein design, where hallucinated structures absent from nature have been experimentally validated and shown to exhibit stable configurations and functional properties in a range of tested scenarios.
  • Robotic navigation, via the "Learning from Hallucinations" paradigm, where patterns emerging from hallucinations are used to optimize path planning.
  • Domain settings that require different balances — the paper notes creative writing may favor higher w₁ while medical applications demand higher w₂, so the IFS weights can be tuned per deployment.

Industry relevance centers on model selection and prompt design: the results show that a model leading on creative tasks may underperform on precision-critical ones, so teams should choose models and constraint levels based on their application rather than on a single aggregate score.

Future Directions

  • Extending beyond structured question-answering. The framework currently focuses on structured tasks and neglects areas such as literary writing; the authors plan to cover broader creativity tasks.
  • Improving IH measurement. The paper identifies potential bias from inductive prompts and calls for refined IH metrics, domain-specific prompts, and human-involved multi-dimensional evaluation to better manage IH–DH dynamics.
  • Broadening the benchmark's reach. Application to multimodal and cross-lingual benchmarks is planned to validate generalizability, alongside examination of the ethical implications of promoting IH.
  • Open ethical question about misuse. The AI Ethics section warns that users must not conflate IH outputs with factual content, since that could lead to misinterpreting IH as reality — an open problem for any downstream deployment.

Target Audience

Researchers working on hallucination detection and mitigation will find the IH/DH distinction directly usable as an alternative to binary factual-consistency evaluation. Evaluation and benchmark designers can reuse the TTCT-based metric matrix, the IFS formulation, and the ten-domain task set. Computational creativity and AI-for-science researchers will care most about the claim that hallucinations can act as a "scientific dream machine." Practitioners choosing models or prompt strategies for open-ended generative products also benefit, since the comparative tables show which strategies raise or suppress each hallucination type. The paper does not report the number of human annotators, the identity of the annotators, or inter-annotator agreement statistics in the content available.

Authors’ abstract

Hallucinations in large language models (LLMs) are commonly regarded as errors to be minimized. However, recent perspectives suggest that some hallucinations may encode creative or epistemically valuable content, a dimension that remains underquantified in current literature. Existing hallucination detection methods primarily focus on factual consistency, struggling to handle heterogeneous scientific tasks and balance creativity with accuracy. To address these challenges, we propose HIC-Bench, a novel evaluation framework that categorizes hallucinations into Intelligent Hallucinations (IH) and Defective Hallucinations (DH), enabling systematic investigation of their interplay in LLM creativity. HIC-Bench features three core characteristics: (1) Structured IH/DH Assessment. using a multi-dimensional metric matrix integrating Torrance Tests of Creative Thinking (TTCT) metrics (Originality, Feasibility, Value) with hallucination-specific dimensions (scientific plausibility, factual deviation); (2) Cross-Domain Applicability. spanning ten scientific domains with open-ended innovation tasks; and (3) Dynamic Prompt Optimization. leveraging the Dynamic Hallucination Prompt (DHP) to guide models toward creative and reliable outputs. The evaluation process employs multiple LLM judges, averaging scores to mitigate bias, with human annotators verifying IH/DH classifications. Experimental results reveal a nonlinear relationship between IH and DH, demonstrating that creativity and correctness can be jointly optimized. These insights position IH as a catalyst for creativity and reveal the ability of LLM hallucinations to drive scientific innovation.Additionally, the HIC-Bench offers a valuable platform for advancing research into the creative intelligence of LLM hallucinations.

Read the original paper