Research
Toward Human-Centered Readability Evaluation
Overview Research area: Natural Language Processing (NLP) for health text simplification, positioned at the intersection of NLP evaluation, Human-Computer Interaction (HCI), and health communication.

- arXiv
- 2510.10801
- Published
- 2025-10-12
- Authors
- Bahar İlgen, Georges Hattab
AI summary
Overview
Research area: Natural Language Processing (NLP) for health text simplification, positioned at the intersection of NLP evaluation, Human-Computer Interaction (HCI), and health communication.
Technical level: Intermediate. The paper is a conceptual position and framework proposal rather than an experimental study; it assumes familiarity with common NLP metrics (BLEU, SARI, FKGL) but requires no implementation background.
Scope: The paper proposes the Human-Centered Readability Score (HCRS), a five-dimensional framework for evaluating simplified health texts using automatic measures combined with structured human feedback, and outlines—but does not yet execute—a protocol for empirical validation.
What This Paper Is About
Text simplification systems rewrite complex health information into easier-to-read versions, but they are typically judged by automatic metrics like BLEU, FKGL, and SARI that only capture surface features such as n-gram overlap, sentence length, and syllables per word. These metrics say nothing about whether a simplified health message is actually clear, trustworthy, respectful, culturally appropriate, or actionable for the people who need it. The authors argue that readability in health contexts is relational and context-sensitive, and they propose a five-dimension evaluation framework that integrates automatic scoring with structured user feedback to close this gap.
Key Contributions
- Redefinition: Reconceptualizing readability for health text simplification from a human-centered perspective, drawing on HCI and health communication research rather than treating readability as a property of text alone.
- Framework: Proposing HCRS, a five-dimension conceptual model—clarity, trustworthiness, tone appropriateness, cultural relevance, and actionability—to guide both evaluation and system design, including hybrid scoring equations that combine automatic features with human Likert ratings.
- Coverage audit: Presenting a table (Table 1) mapping common automatic metrics (BLEU, SARI, FKGL, BERTScore, QuestEval, SALSA) against these user-centered dimensions, showing that no widely used metric comprehensively covers clarity, trustworthiness, and actionability.
- Agenda: Outlining a participatory, human-in-the-loop evaluation protocol and a future research agenda bridging NLP and HCI, including how HCRS could act as a complementary evaluation layer and calibration signal for Reinforcement Learning with AI Feedback (RLAIF) pipelines in sensitive domains.
Main Findings
-
Surface metrics miss user-centered qualities: BLEU, SARI, and FKGL address lexical and (partially) syntactic simplicity but do not address trustworthiness or actionability; BERTScore and QuestEval address semantic adequacy only; SALSA covers lexical and structural edits but does not assess whether outputs are clear to diverse users.
-
Weak correlation with human judgment: The paper cites Alva-Manchego et al. (2021), who found that commonly used automatic metrics such as BLEU and SARI typically show only low-to-moderate correlation with human judgments on simplification quality, particularly when multiple rewriting operations are involved. In that systematic comparison, BERTScore Precision achieved the highest overall correlation, but performance dropped sharply for high-quality outputs.
-
Simplification can improve comprehension: Leroy et al. (2022) found that simplified health texts significantly improved comprehension accuracy, boosting correct recall from 33% to 59%, with the largest gains among participants with lower education or limited English proficiency.
-
Simplicity does not guarantee effective communication: The paper notes that even high-SARI outputs can be perceived as emotionally flat or insufficiently actionable, and that surface-level gains in simplicity do not necessarily ensure health information is perceived as accessible or actionable.
-
Combining metrics helps: Choi et al. (2024) is cited as showing that dynamically combining multiple evaluation metrics (lexical and semantic) yields much stronger alignment with human ratings than any single metric.
-
RLAIF scaling versus alignment: RLAIF reduces the cost and time of annotation by more than an order of magnitude while achieving competitive—and sometimes superior—results compared to RLHF on standard benchmarks such as win rate and harmlessness, but its optimization targets remain generic (helpfulness, harmlessness) and do not inherently capture trustworthiness, cultural relevance, or actionability.
-
No empirical validation reported: The paper explicitly states that HCRS has not been empirically validated on large-scale, diverse user populations, that dimension weightings are conceptual and require calibration, and that no pilot study was conducted within the scope of this work.
Methodology in Plain English
The authors did not run experiments. Instead, they reviewed existing automatic evaluation metrics used for text simplification and mapped what each one actually measures against five qualities they argue matter for health communication. They then constructed a scoring framework in which each of the five dimensions is estimated by blending automatic computational signals with structured human ratings.
For clarity, the automatic side uses readability indices (FKGL, SMOG), jargon detection, and cohesion tools such as Coh-Metrix, combined with user-rated comprehension and ease-of-reading survey items on a 5-point Likert scale or direct comprehension quizzes. For trustworthiness, the framework detects explicit source attribution, institutional language, transparency features, and domain authority markers, supplemented by participant ratings of credibility, transparency, and author reliability. Tone appropriateness combines automatic pragmatic feature extraction (politeness classifiers such as the Stanford Politeness Classifier, formality indices, indicators like indirectness, mitigation, and hedging), sentiment and emotion analysis using transformer models (BERT, RoBERTa), empathy and support classifiers (e.g., EmpathBERT), and lexical diversity and intensity measures (intensifiers, modals, evidentials, negative polarity items), blended with human Likert ratings in a weighted equation. Cultural relevance combines automatic entity matching via named entity recognition, idiom and cultural expression matching, and multilingual embedding similarity, plus human ratings of familiarity and inclusivity. Actionability combines automatic detection of directive and imperative language, procedural and instruction cues, and action-associated entities (temporal, agent, location), plus human ratings of perceived next-step clarity. In each case the weights are to be calibrated empirically on validation data to maximize alignment with user perceptions.
For the human side, the paper proposes participatory evaluation: lightweight annotation interfaces where end users rate sentences on the five dimensions, micro-surveys lasting 5–10 minutes, and stakeholder workshops with patients, clinicians, or domain experts that feed back into HCRS weighting—an iterative loop of micro-ratings, review, and refinement.
Why This Matters
Impact on research: The paper challenges the assumption that surface-level scores such as BLEU, SARI, and FKGL are adequate proxies for readability in high-stakes domains, and it proposes testable research questions (Q1–Q3) about which metrics align with human-centered dimensions, whether a composite score beats the best standalone metric, and how HCI techniques can be embedded into evaluation pipelines. It also positions HCRS as a complementary layer for RLAIF, arguing that synthetic evaluators inherit generic preferences from pre-trained LLMs and risk amplifying existing biases.
Real-world applications:
- Public health communication: medication instructions, risk explanations, and care recommendations that must be understandable to people with limited health literacy.
- Immunization and vaccine information: the paper uses a vaccine information use-case as a proof-of-concept, where patients rate clarity and trustworthiness via inline sliders and health professionals recalibrate actionability guidelines.
- Clinical and patient-facing materials where tone, respect, and avoidance of blame affect whether readers feel addressed rather than alienated.
- Culturally tailored health messaging for multilingual and marginalized audiences, where cultural loss during simplification can reduce both accessibility and trust.
Industry relevance: Organizations deploying text simplification or LLM-based health content generation need evaluation signals that predict user comprehension and trust, not just n-gram overlap. The framework offers a route to calibrating models with structured, interpretable human feedback channels (for example, "Too technical", "Missing information", "Poorly structured") rather than binary ratings, and it warns against over-reliance on RLAIF in sensitive domains without human-in-the-loop validation.
Future Directions
- Empirical validation: Conduct small-scale pilot studies using micro-surveys and participatory workshops on real health communication materials, since the framework has not yet been tested on large-scale, diverse user populations.
- Weight calibration: Move dimension weightings from conceptual to empirically calibrated against real-world user judgments.
- Domain and language generalization: Test HCRS beyond health communication, and retune sociocultural and emotional dimensions—along with the language resources and annotation protocols they depend on—for new target populations, noting that concepts such as trustworthiness and actionability may need redefinition per domain.
- Practical integration into NLP pipelines: Combine automatic readability features with lightweight human feedback modules, test participatory feedback pipelines at scale, explore adaptive dimension weighting, and examine how HCRS can serve as a calibration signal for RLAIF-based training.
Target Audience
This paper is most useful for NLP researchers and practitioners working on text simplification and health-focused language technologies, HCI and participatory design researchers interested in evaluation methodology, and health communication specialists, public health agencies, and clinical content teams who need to judge whether simplified materials are actually understood, trusted, respected, and acted upon. It is also relevant to researchers working on model alignment and RLHF/RLAIF who need to understand where synthetic feedback falls short in high-stakes domains.
Authors’ abstract
Text simplification is essential for making public health information accessible to diverse populations, including those with limited health literacy. However, commonly used evaluation metrics in Natural Language Processing (NLP), such as BLEU, FKGL, and SARI, mainly capture surface-level features and fail to account for human-centered qualities like clarity, trustworthiness, tone, cultural relevance, and actionability. This limitation is particularly critical in high-stakes health contexts, where communication must be not only simple but also usable, respectful, and trustworthy. To address this gap, we propose the Human-Centered Readability Score (HCRS), a five-dimensional evaluation framework grounded in Human-Computer Interaction (HCI) and health communication research. HCRS integrates automatic measures with structured human feedback to capture the relational and contextual aspects of readability. We outline the framework, discuss its integration into participatory evaluation workflows, and present a protocol for empirical validation. This work aims to advance the evaluation of health text simplification beyond surface metrics, enabling NLP systems that align more closely with diverse users' needs, expectations, and lived experiences.