Research
CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
Overview Research area: Natural Language Processing, specifically large language model (LLM) hallucination mitigation and prompt refinement. Technical level: Intermediate. The paper assumes familiarit

- arXiv
- 2510.12029
- Published
- 2025-10-14
- Authors
- Jung-Woo Shim, Yeong-Joon Ju, Ji-Hoon Park, Seong-Whan Lee
AI summary
Overview
- Research area: Natural Language Processing, specifically large language model (LLM) hallucination mitigation and prompt refinement.
- Technical level: Intermediate. The paper assumes familiarity with instruction fine-tuning, LoRA, perplexity, and standard generation metrics (BLEU, ROUGE, METEOR), but its core idea is straightforward.
- Scope: The paper proposes and evaluates Curative Prompt Refinement (CPR), a plug-and-play framework that uses a fine-tuned small language model (SLM) to clean ill-formed user prompts and append generated informative task descriptions before an inference LLM responds.
What This Paper Is About
LLMs sometimes produce plausible but factually incorrect content ("hallucinations"), and the authors argue that a frequently overlooked cause is the user's own prompt: vague, grammatically broken, or underspecified inputs that force the model to guess at intent. Rather than changing the LLM itself or correcting its output afterward, CPR repairs the prompt before inference, using a small fine-tuned model to fix errors and add contextual descriptions so the downstream LLM's actual task is unambiguous.
Key Contributions
- A plug-and-play prompt-side framework for hallucination mitigation. CPR is model-agnostic, lightweight, and requires no external knowledge base, no reinforcement learning, and no access to large-scale compute; it sits in front of any inference LLM.
- A three-task fine-tuning dataset for an SLM. The authors build instruction fine-tuning pairs from three sources: the Wikipedia English dataset (WikiEn) for punctuation/grammar correction, the Multi-domain Question Rewriting dataset (MQR) for paraphrase and word substitution, and the Wikidata Description dataset (WikiD) for description generation.
- A description generation and reranking algorithm. Multiple candidate descriptions are generated for a cleaned prompt and then filtered and reranked by perplexity, with generation halted when perplexity crosses a threshold, so only the most coherent descriptions are kept (Algorithm 1).
- Empirical evidence across many SLMs and inference models. CPR is tested with fine-tuned SLMs including Gemma (2B), Phi-2 (2.7B), Llama-2 (7B), Phi-3 (3.8B), Qwen (1.5) (4B), Zephyr (3B), and FLAN-T5 XL (2.85B), paired with Llama-2 (7B) and GPT-3.5 as inference models.
Main Findings
- Refined prompts reduce measured hallucination and raise content quality. On the original (ill-formed) prompts, Llama-2 (7B) as inference model scored HI 0.51 and CQS 0.38, while GPT-3.5 scored HI 0.16 and CQS 0.51. With CPR using Llama-2 (7B) as the fine-tuned SLM, Llama-2 inference reached HI 0.13, CQS 0.75, WR 0.92, and GPT-3.5 inference reached HI 0.04, CQS 0.83, WR 0.96.
- Even small SLMs work. With CPR using Gemma (2B) as the refiner and GPT-3.5 as inference model, the results were HI 0.07, CQS 0.68, and WR 0.92. The paper reports no significant difference in hallucination reduction between a 2B and a 7B refiner, though the larger refiner improved quality and win rate.
- The descriptions matter. CPR without descriptions performs worse than full CPR; the authors call the absence of descriptions a marked reduction in effectiveness. Several CPR-with-description rows in Table I show CQS values of 0.67–0.83 versus 0.59–0.75 for CPR without descriptions.
- Gains are larger for smaller inference models. The paper reports that the degree of improvement in inference models diminishes as model size increases, with smaller models showing more substantial improvements.
- Prompt cleaning and paraphrasing improve after fine-tuning. BLEU, ROUGE, and METEOR all rise: Gemma went from 11.2/45.3/31.3 to 21.1/54.2/32.1; Phi-2 from 13.1/46.6/31.7 to 21.7/56.3/35.6; Llama-2 from 12.8/48.1/31.5 to 23.1/56.2/36.5.
- Generated descriptions become more relevant and coherent after fine-tuning. Llama-2 improved from relevance 63.4 and coherence 62.5 to 80.5 and 82.1; Phi-2 from 61.4/58.8 to 73.4/81.4; Gemma from 61.1/56.8 to 71.1/79.5. The paper reports 0 for the fine-tuned entries of Phi-3.5 mini (3B) and Llama-3.2 (3.21B) in Table III, and 0 for their GPT-3.5 columns and win rates in Table I.
- CPR and a post-processing method are complementary. Against SelfCheckGPT, on prompts of low ill-formedness (score above 0.2), CPR with Gemma (2B) scored HI 0.23, CQS 0.42, WR 0.61 and CPR with Llama-2 (7B) scored HI 0.21, CQS 0.62, WR 0.68, while SelfCheckGPT scored HI 0.19, CQS 0.58, WR 0.71. On highly ill-formed prompts (score under 0.2, "High"), SelfCheckGPT scored HI 0.37, CQS 0.42, WR 0.51 versus CPR with Llama-2 (7B) at HI 0.23, CQS 0.57, WR 0.69. Combining both gave the best numbers: HI 0.03, CQS 0.91, WR 0.98 on the Low set and HI 0.05, CQS 0.88, WR 0.99 on the High set.
- Headline win rates. The abstract reports over a 90% win rate over original prompts with no external knowledge; the introduction states a 96% win rate over original ill-formed prompts with GPT-3.5 as the inference model, and a 99% win rate over highly ill-formed prompts when a post-processing hallucination mitigation approach is applied.
Methodology in Plain English
The authors leave the LLM alone and fix the input instead. First they assemble a training set from three public sources and turn it into instruction-style prompt/completion pairs: roughly 10,000 Wikipedia English entries (the GPT-3.5 API verified and refined these to be error-free, and the same API introduced punctuation errors into copies to create paired examples), 2,114 paraphrased question pairs from MQR, and 10,000 keyword-and-short-description pairs from Wikidata selected by lookup frequency. A small language model is then fine-tuned on these pairs using instruction fine-tuning plus LoRA, which updates only a small subset of parameters to avoid catastrophic forgetting.
At inference time the fine-tuned SLM does two things. It cleans and paraphrases the user's ill-formed prompt using the grammar skills learned from the Wikipedia data and the rewriting skills learned from MQR. Then, because a clean prompt can still be too thin on context, it generates several candidate descriptions of the task, monitors perplexity during generation, and stops when perplexity reaches the predefined threshold (reported as 15 in the method section). The top-k descriptions by lowest resting perplexity are reranked and selected, and the cleaned prompt plus these selected descriptions are passed to the inference LLM.
Evaluation used 8,000 user queries drawn from the Google Well-formed Query dataset, each with a score below 0.5, anonymized and randomized. Fine-tuning ran on a single NVIDIA RTX A6000 GPU and SLM inference on a single NVIDIA TITAN V GPU. Three metrics judged by the GPT-3 API were used: Hallucination Index (HI, 0 to 1, lower is better), Content Quality Score (CQS, 0 to 1, higher is better), and Win Rate (WR). Prompt refinement was measured with BLEU, ROUGE, and METEOR; description quality was measured with relevance and coherence, also judged by the GPT-3 API, with up to five descriptions generated per query and generation halted at a perplexity threshold of 5 in that experiment.
Why This Matters
- Impact on research: The work reframes hallucination as partly an input-quality problem rather than purely a model-internals or post-hoc verification problem, and it argues that prior prompt-refinement work relies on large models, human intervention, reinforcement learning, or external knowledge bases, which the authors position as costly and hard to scale. It also offers evidence that an SLM, not an LLM, is sufficient for the refinement step.
- Real-world applications:
- Consumer chatbots and search assistants where non-expert users type vague or ungrammatical queries.
- Customer support and internal help desks that route raw user messages into a downstream LLM.
- Educational or reference tools where factual accuracy of the response matters more than verbosity.
- Low-resource deployment settings, since the framework avoids external knowledge bases and high-end compute.
- Industry relevance: CPR is designed as a plug-and-play preprocessing layer that works across LLM architectures without tailored adjustments, which matters for organizations that want to improve output reliability without retraining or swapping their existing model. The finding that CPR and a post-processing method such as SelfCheckGPT combine for the best scores suggests it can sit alongside, rather than replace, current hallucination-mitigation infrastructure.
Future Directions
- Build a higher-quality fine-tuning dataset. The authors state that their dataset was manually crafted, and note that a better-crafted dataset could amplify CPR's effectiveness.
- Test at larger model scales and wider model ranges. They report that resource constraints limited their evaluation scope, restricting a thorough exploration of scalability and performance across larger models.
- Investigate the CPR plus post-processing combination further. Since CPR outperformed SelfCheckGPT on highly ill-formed prompts while SelfCheckGPT was more effective on minimally flawed prompts, and the combination scored best on every metric (HI 0.03/CQS 0.91/WR 0.98 on the Low set; HI 0.05/CQS 0.88/WR 0.99 on the High set), the boundary between the two regimes remains an open question.
- Clarify why description generation did not converge for some models. The paper reports 0 values for Phi-3.5 mini (3B) and Llama-3.2 (3.21B) after fine-tuning in Tables I and III, which the paper does not explain in the provided content.
Target Audience
Readers who will benefit most are NLP researchers and graduate students working on hallucination mitigation, prompt optimization, or parameter-efficient fine-tuning; machine learning engineers who need a lightweight, model-agnostic preprocessing layer for production LLM systems; and practitioners without access to large-scale compute who want to improve LLM output reliability using only small models. The paper is also useful for anyone weighing prompt-side interventions against model-side or post-hoc correction methods.
Authors’ abstract
Recent advancements in large language models (LLMs) highlight their fluency in generating responses to diverse prompts. However, these models sometimes generate plausible yet incorrect ``hallucinated" facts, undermining trust. A frequent but often overlooked cause of such errors is the use of poorly structured or vague prompts by users, leading LLMs to base responses on assumed rather than actual intentions. To mitigate hallucinations induced by these ill-formed prompts, we introduce Curative Prompt Refinement (CPR), a plug-and-play framework for curative prompt refinement that 1) cleans ill-formed prompts, and 2) generates additional informative task descriptions to align the intention of the user and the prompt using a fine-tuned small language model. When applied to language models, we discover that CPR significantly increases the quality of generation while also mitigating hallucination. Empirical studies show that prompts with CPR applied achieves over a 90\% win rate over the original prompts without any external knowledge.