Research
Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs
Overview Research area: Natural Language Processing — abstractive summarization of consumer health questions (CHQs), with a focus on faithfulness (factual consistency with the source) rather than flue

- arXiv
- 2511.10768
- Published
- 2025-11-13
- Authors
- Ajwad Abrar, Nafisa Tabassum Oeshy, Prianka Maheru, Farzana Tabassum, Tareque Mohmud Chowdhury
AI summary
Overview
Research area: Natural Language Processing — abstractive summarization of consumer health questions (CHQs), with a focus on faithfulness (factual consistency with the source) rather than fluency alone, evaluated cross-lingually in English and Bangla.
Technical level: Intermediate. Readers should be comfortable with standard summarization concepts (extractive vs. abstractive methods, ROUGE, BERTScore) and with the idea of fine-tuning a large language model.
Scope: The paper proposes a framework combining TextRank-based sentence extraction and medical named entity recognition (NER) with a fine-tuned LLaMA-2-7B model, and evaluates it on MeQSum (English) and BanglaCHQ-Summ (Bangla) using both quality and faithfulness metrics plus a human evaluation.
What This Paper Is About
Consumer health questions posted online are often long, verbose, and redundant, making it hard for clinicians to quickly identify the core concern. Automatic summarization could help, but in medicine a summary that distorts a symptom, medication, or relationship is dangerous — and standard metrics like ROUGE or BERTScore do not measure whether a summary stays factually faithful to the source. The paper's goal is a summarization framework that produces concise, readable CHQ summaries while explicitly preserving medically important content, tested in both a high-resource (English) and a low-resource (Bangla) setting.
Key Contributions
- Hybrid extractive–abstractive design. The framework integrates TextRank-based sentence extraction with abstractive generation by an LLM, aiming to improve both informativeness and reliability of the resulting summaries.
- Medical NER as a faithfulness constraint. Medical named entity recognition is used to ensure that critical entities are identified and preserved in the summaries before the LLM generates text.
- First cross-lingual faithfulness evaluation for CHQ summarization. The authors report what they describe as the first cross-lingual evaluation of faithfulness in medical CHQ summarization, covering English (MeQSum) and Bangla (BanglaCHQ-Summ).
- Improvements over zero-shot baselines and prior systems. The approach is reported to outperform zero-shot baselines and prior state-of-the-art systems on the evaluated metrics, with human verification of faithfulness.
Main Findings
-
Zero-shot LLM performance is weak. On MeQSum, zero-shot LLaMA-2-7B without fine-tuning reached ROUGE-1 of 21.97, ROUGE-2 of 6.48, ROUGE-L of 19.98, BERTScore of 0.60, readability (Flesch Reading Ease) of 65.16, SummaC of 0.28, and AlignScore of 21.80.
-
Fine-tuning sharply improves both quality and faithfulness. Fine-tuning without TextRank raised MeQSum scores to ROUGE-1 44.23, ROUGE-2 27.36, ROUGE-L 41.55, BERTScore 0.71, readability 70.21, SummaC 0.31, and AlignScore 38.45.
-
Adding TextRank yields further gains. The FT + TextRank setting produced ROUGE-1 47.07, ROUGE-2 29.44, ROUGE-L 44.08, BERTScore 0.72, readability 70.69, SummaC 0.37, and AlignScore 45.65 on MeQSum.
-
Best-of-3 selection gives the strongest results, and the choice of selector matters. Selecting the best of three candidates by ROUGE-1 gave the top quality scores on MeQSum (ROUGE-1 50.50, ROUGE-2 34.38, ROUGE-L 47.74, BERTScore 0.74, readability 71.56, SummaC 0.40, AlignScore 39.24), while selecting by SummaC gave the highest factual consistency (SummaC 0.57) with ROUGE-1 48.27, ROUGE-2 31.38, ROUGE-L 45.34, BERTScore 0.73, readability 71.56, and AlignScore 45.91.
-
A temperature trade-off exists. Sweeping t over {0.1, 0.3, 0.5, 0.7, 0.9} showed that lower t favors ROUGE while higher t favors SummaC; the authors adopted t = 0.7 as a balanced choice giving peak faithfulness with competitive ROUGE.
-
The method compares favorably with prior systems on MeQSum. Against Mixtral-8x7B-Inst. (ROUGE-1 32.47, ROUGE-2 36.38, ROUGE-L 16.86, BERTScore 0.72; readability, SummaC and AlignScore not reported) and BioBART + FaMeSumm (ROUGE-1 31.76, ROUGE-2 11.71, ROUGE-L 29.64, BERTScore 0.74, SummaC 0.46; readability and AlignScore not reported), the authors' best-of-3 method reported ROUGE-1 50.50, ROUGE-2 34.38, ROUGE-L 47.74, BERTScore 0.74, readability 71.56, SummaC 0.57, and AlignScore 0.46.
-
The pattern holds in Bangla. On BanglaCHQ-Summ, zero-shot scored ROUGE-1 19.10, ROUGE-2 8.21, ROUGE-L 18.97, BERTScore 0.62, SummaC 0.22; fine-tuning without TextRank scored 28.24 / 14.22 / 24.54 / 0.71 / 0.26; FT + TextRank scored 30.71 / 15.71 / 28.95 / 0.74 / 0.28; best-of-3 by ROUGE-1 scored 32.35 / 16.32 / 29.09 / 0.76 / 0.29; best-of-3 by SummaC scored 30.92 / 15.74 / 27.35 / 0.73 / 0.32. Readability and AlignScore are not reported for Bangla.
-
Human evaluation supports faithfulness. A medical doctor evaluated MeQSum summaries on two binary (yes/no) questions — whether all critical information from the source was retained, and whether the summary was factually consistent — with a summary counted as faithful only if both conditions were met. 82% of summaries satisfied both criteria; the abstract describes this as over 80%.
-
Stated limitation. Experiments used a single LLM (LLaMA-2-7B) and two languages.
Methodology in Plain English
The researchers start with two benchmark collections of consumer health questions paired with reference summaries: MeQSum, with 1,000 expert-validated English question–summary pairs, and BanglaCHQ-Summ, with 2,350 annotated Bangla pairs.
Preprocessing standardizes both datasets into "question" and "summary" fields and identifies overlapping medical entities and negation terms, so that important details — including things a patient denies having — are not lost. Medical named entity recognition supplies the list of medically meaningful terms. The TextRank algorithm is then applied to pick out the sentences in each question that contain those medical entities and query-related words, filtering noise and keeping source-grounded content before any text is generated. This extraction step acts as a guardrail: the generative model works from material that is already known to be medically relevant.
The generative component is LLaMA-2-7B, fine-tuned with low-rank adaptation on each dataset. The authors compare four settings: zero-shot (no fine-tuning), fine-tuning without TextRank, fine-tuning with TextRank-selected sentences, and a "best-of-3" setup in which three candidate summaries are generated and one is chosen either by ROUGE-1 (favoring surface quality) or by SummaC (favoring factual consistency).
Evaluation uses two families of metrics. General quality is measured by ROUGE-1, ROUGE-2, ROUGE-L and BERTScore, with Flesch Reading Ease for readability. Faithfulness is measured by SummaC and AlignScore, which score factual consistency between summary and source. Finally, a medical doctor performed a binary human evaluation on MeQSum summaries.
Why This Matters
Impact on research. The paper argues that faithfulness is an underexplored dimension compared to readability or general accuracy in medical summarization, and that most existing methods do not explicitly address it. By pairing standard quality metrics with SummaC and AlignScore and reporting a cross-lingual comparison, it pushes evaluation practice toward factual consistency, and it demonstrates that a low-resource language (Bangla) can be handled with the same framework.
Real-world applications:
- Triage of online patient queries, condensing verbose consumer health questions so clinicians can reach the core concern faster.
- Patient-facing health portals that restate a user's question or a clinical answer in a shorter, readable form without altering medical meaning.
- Cross-lingual health information access, particularly for Bangla-speaking users where large-scale annotated CHQ resources are scarce.
- Screening and quality assurance pipelines that flag summaries whose medical entities or negations do not match the source.
Industry relevance. The method's extractive pre-filter plus fine-tuned open model is a comparatively cheap recipe that does not require a large proprietary model, and best-of-3 selection with a faithfulness scorer is a directly implementable inference-time control. The temperature finding gives practitioners an explicit quality-versus-faithfulness dial. For any organization deploying LLMs in a health context, the reported 82% human-judged faithfulness rate indicates both the promise and the remaining risk.
Future Directions
- Broaden beyond one model. The authors note that experiments used only LLaMA-2-7B and call for exploring multiple LLMs.
- Extend to more languages. The current work covers two languages, and the authors propose adapting the framework to broader multilingual settings.
- Few-shot settings. Adapting the framework to few-shot scenarios is listed as a future direction, which would reduce dependence on large annotated datasets such as the 1,000-pair MeQSum and 2,350-pair BanglaCHQ-Summ.
- Transfer to other sensitive domains and expert-in-the-loop use. The authors suggest extending to legal and financial texts and incorporating expert feedback, with the aim of improving robustness and integration into clinical workflows.
Target Audience
This paper is most useful to NLP researchers working on summarization and factual-consistency evaluation, and to researchers specifically focused on biomedical and clinical text processing who need a cross-lingual reference point. It also suits applied machine learning engineers and health-tech product teams evaluating whether a fine-tuned open LLM with extractive grounding can be trusted for patient-facing summarization. Because the framework is described step by step and evaluated with widely used metrics, readers with intermediate NLP background — including graduate students beginning work in medical NLP — can follow it without specialized clinical training.
Authors’ abstract
Summarizing consumer health questions (CHQs) can ease communication in healthcare, but unfaithful summaries that misrepresent medical details pose serious risks. We propose a framework that combines TextRank-based sentence extraction and medical named entity recognition with large language models (LLMs) to enhance faithfulness in medical text summarization. In our experiments, we fine-tuned the LLaMA-2-7B model on the MeQSum (English) and BanglaCHQ-Summ (Bangla) datasets, achieving consistent improvements across quality (ROUGE, BERTScore, readability) and faithfulness (SummaC, AlignScore) metrics, and outperforming zero-shot baselines and prior systems. Human evaluation further shows that over 80\% of generated summaries preserve critical medical information. These results highlight faithfulness as an essential dimension for reliable medical summarization and demonstrate the potential of our approach for safer deployment of LLMs in healthcare contexts.