Research
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
Overview Research area: Natural Language Processing — specifically the interaction between hallucination-reduction techniques and creative generation in large language models, with an eye toward AI-as

- arXiv
- 2512.11509
- Published
- 2025-12-12
- Authors
- Mohor Banerjee, Nadya Yuki Wangsajaya, Syed Ali Redha Alsagoff, Min Sen Tan, Zachary Choy Kit Chun, Alvin Chan Guo Wei
AI summary
Overview
Research area: Natural Language Processing — specifically the interaction between hallucination-reduction techniques and creative generation in large language models, with an eye toward AI-assisted scientific discovery.
Technical level: Intermediate. The paper assumes familiarity with LLM decoding, retrieval augmentation, and benchmark-based evaluation, but its central question is stated in plain terms and its main findings are readable without deep technical background.
Scope: The paper empirically measures how three hallucination-reduction methods (Chain of Verification, Decoding by Contrasting Layers, and Retrieval-Augmented Generation) change convergent and divergent creativity across multiple model families and scales on two creativity benchmarks.
What This Paper Is About
Large language models hallucinate, producing factually incorrect content, and many methods exist to suppress this behavior. What has not been studied is whether those same methods damage a model's creative ability, which is a problem for scientific work that needs both factual accuracy and novel hypothesis generation. The authors apply three hallucination-reduction techniques to a range of models and measure both convergent creativity (solving correctly within constraints) and divergent creativity (generating varied, novel ideas).
Key Contributions
- A systematic demonstration that hallucination-reduction methods affect divergent creativity in opposite directions while leaving convergent creativity largely unchanged — CoVe enhances it, DoLa suppresses it, and RAG has minimal impact.
- Evidence that this pattern generalizes across model families (LLaMA, Qwen, Mistral) and scales (1B to 70B parameters), suggesting it is not an artifact of model size or architecture.
- A mechanistic investigation using linear probes showing that early transformer layers correlate more strongly with creativity than later layers, offering an explanation for why DoLa's contrastive decoding suppresses creative output.
- A preliminary method that amplifies creativity-correlated layers and suppresses anti-correlated layers, which improved divergent creativity in LLaMA and Qwen-coder without compromising convergent creativity.
Main Findings
-
CoVe increases divergent creativity: On NeoCoder, LLaMA 1B achieves the highest improvement, peaking at around 12.5% above baseline, while LLaMA 8B and Qwen-coder 7B show moderate gains of 2–4%. Mistral 7B deviates from the trend, ranging from –3% to +2%. On CS4, LLaMA 8B and LLaMA 1B maintain steady improvements of approximately 5–8%, and Mistral 7B shows a modest but stable increase of around 2%.
-
DoLa reduces divergent creativity: On NeoCoder, LLaMA 1B ranges from approximately –2.5% to –1%, and Mistral 7B remains mostly below baseline, reaching around –2.5% with a single rise to about +3% at the third state. LLaMA 8B, Qwen-coder 7B, and LLaMA 70B remain nearly constant, between –1% and –0.5%. On CS4 the reduction is more pronounced: LLaMA 1B drops to around –8%, while LLaMA 8B and Mistral 7B show smaller decreases of roughly –3% to –2%.
-
RAG has no meaningful effect on divergent creativity: Across NeoCoder evaluations, LLaMA 70B shows small positive shifts up to about +3%, LLaMA 8B declines to around –5% with no positive deviation, Qwen-coder 7B moves between roughly –0.5% and +1.5%, LLaMA 1B varies between approximately –3% and +1.5%, and Mistral 7B ranges from about –2.5% to +2.5%. RAG was applied only to NeoCoder, since CS4's open-ended story generation lacks a defined retrieval corpus.
-
Convergent creativity is largely unaffected: Performance differences relative to baseline for CoVe, RAG, and DoLa remain minimal across both datasets and all models, fluctuating around zero without a consistent trend.
-
Early layers track creativity: Linear probes trained to predict whether the model will generate a divergently creative output showed that the top 5 creativity-correlated layers often cluster at early layers. This supports the hypothesis that DoLa's contrastive operation removes layer activations responsible for creative generation.
-
Layer modulation can boost creativity: Amplifying the top 5 creativity-correlated layers while suppressing the bottom 5 anti-correlated layers boosted divergent creativity in LLaMA and Qwen-coder without compromising convergent creativity. CS4 results for this method were evaluated on LLaMA 8B only, due to computation constraints.
-
Creative probes are model-specific: Probes trained only on LLaMA 3.1 8B generations and tested on Qwen-coder and Mistral yielded poor results. Replacing the training data with generations from those specific models produced substantial improvements, which the authors attribute to distribution shift in activation space.
-
The authors' hypothesis was wrong: They expected hallucination-reduction techniques would generally suppress creativity. Instead, the methods produced opposing effects on divergent creativity.
Methodology in Plain English
The researchers took language models and generated outputs in two settings, once with no intervention (the baseline) and once with each of three hallucination-reduction methods switched on. They then compared the results as percentage changes from baseline using the formula (method score − baseline score) / baseline score × 100, and reported the mean of three independent runs per configuration.
The three methods work differently. Chain of Verification has the model draft an answer, generate verification questions about that draft, answer them, and produce a refined final response. Decoding by Contrasting Layers compares predictions from a higher layer against an earlier, "premature" layer chosen dynamically at each step by finding which layer's output differs most from the final layer using Jensen-Shannon divergence, then subtracts the earlier layer's logits from the later layer's. Retrieval-Augmented Generation retrieves relevant external documents before generating, here using ColBERTv2 following the RAGLAB framework, indexing a corpus of coding tutorials, library documentation, GitHub repositories, and programming solutions from CodeRAG-bench, retrieving the top 3 ranked segments via cosine similarity and appending them to the prompt.
Creativity was measured with two benchmarks. NeoCoder uses 199 CodeForces problems, each paired with around 30 human-written correct solutions, and a maximum of 5 constraints (T = 5). Convergent creativity measures the proportion of correct solutions that also satisfy all constraints; divergent creativity measures the proportion of atomic programming techniques used that were not observed in human-written solutions. CS4 expands each of 50 instructions into 39 cumulative constraints, segmented into sets of 7, 15, 23, 31, and 39, giving 250 unique prompts. It measures constraint satisfaction, coherence via pairwise LLM-as-a-Judge comparisons against a baseline story at 23 constraints rated 1–5 and normalized to [0,1], diversity via Dist-n (the product of unique n-gram ratios for n = 2 to 4), and a composite creativity score QUC (Quality Under n Constraints) equal to normalized coherence multiplied by constraint satisfaction. GPT-5-mini was used as the LLM-as-a-Judge throughout.
For the probing work, the authors trained linear probes, inspired by Inference-Time Intervention, to predict whether a divergently creative output would be produced. They curated only convergently creative answers for training and augmented the dataset by using outputs from hallucination-reduction methods as conditioning context, finding through validation on a small NeoCoder subset that using about 40% of the output as an input signal worked best. Since DoLa operates on whole layers rather than individual attention heads, attention head-level correlations were aggregated into layer-level scores.
Why This Matters
The paper reframes hallucination and creativity as a trade-off that researchers can consciously navigate rather than a single knob to turn down. For AI-assisted scientific discovery, which needs both factual reliability and hypothesis generation, knowing that CoVe preserves or improves divergent thinking while DoLa dampens it changes which method a practitioner would pick. The dissociation between convergent and divergent creativity — and the demonstration that one can be improved without degrading the other — suggests the two capabilities are not locked together.
Real-world applications implied by the work:
- Scientific hypothesis generation: Selecting CoVe-style verification when a system must propose unconventional research directions without sacrificing factual grounding.
- Factual question answering and knowledge work: Recognizing that DoLa-style contrastive decoding, suited to factuality-critical deployments, carries a cost in idea diversity.
- Creative writing and story generation: The CS4 results, where CoVe improved divergent creativity by roughly 5–8% for LLaMA models, inform tooling for open-ended content generation.
- Competitive programming assistance: NeoCoder results describe how novelty in solution technique changes under factual-control methods, relevant to code generation and solutions that go beyond human-written patterns.
Industry relevance: Teams deploying retrieval-augmented or contrastive-decoding systems for factual accuracy have a concrete reason to check creative output as well. The finding that RAG produced model-specific fluctuations in both directions, with LLaMA 8B dropping to around –5%, also indicates that retrieval pipelines need their retrieval quality assessed rather than assumed.
Future Directions
- Develop creativity evaluation frameworks specifically for scientific hypotheses, since programming problems and story generation are only proxies for authentic scientific discovery.
- Conduct ablation studies to isolate the mechanism by which CoVe's questioning process improves divergent creativity, which the authors did not do.
- Systematically measure retrieval quality and test retrieval strategies beyond cosine similarity to determine whether better retrieval could make RAG boost divergent creativity rather than leave it unchanged.
- Determine whether targeted layer amplification scales beyond LLaMA and Qwen-coder and beyond the single-model CS4 evaluation, and whether probes can be made transferable rather than requiring model-specific training.
Target Audience
Researchers and engineers working on LLM factuality, decoding strategies, or retrieval augmentation; AI4Science practitioners who need to choose a hallucination-reduction method for hypothesis-generation pipelines; and anyone studying creativity evaluation in language models, including those interested in interpretability work on how internal representations relate to creative output.
Authors’ abstract
Large Language Models (LLMs) exhibit remarkable capabilities in natural language understanding and reasoning, but suffer from hallucination: the generation of factually incorrect content. While numerous methods have been developed to reduce hallucinations, their impact on creative generations remains unexplored. This gap is particularly critical for AI-assisted scientific discovery, which requires both factual accuracy and creative hypothesis generation. We investigate how three hallucination-reduction techniques: Chain of Verification (CoVe), Decoding by Contrasting Layers (DoLa), and Retrieval-Augmented Generation (RAG), affect creativity in LLMs. Evaluating multiple model families (LLaMA, Qwen, Mistral) at varying scales (1B - 70B parameters) on two creativity benchmarks (NeoCoder and CS4), we find that these methods have opposing effects on divergent creativity. CoVe enhances divergent thinking, DoLa suppresses it, and RAG shows minimal impact. Our findings provide guidance for selecting appropriate hallucination-reduction methods in scientific applications, where the balance between factual accuracy and creative exploration is crucial.