Research
Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?
Overview Research area: Natural language processing / LLM interpretability and trustworthiness, specifically the faithfulness of model-generated explanations in sensitive domains (healthcare and socia

- arXiv
- 2510.24236
- Published
- 2025-10-28
- Authors
- Teague McMillan, Gabriele Dominici, Martin Gjoreski, Marc Langheinrich
AI summary
Overview
- Research area: Natural language processing / LLM interpretability and trustworthiness, specifically the faithfulness of model-generated explanations in sensitive domains (healthcare and social-bias settings).
- Technical level: Intermediate. The paper assumes familiarity with few-shot prompting, Chain-of-Thought, RLHF, and correlation-based evaluation, but explains its causal framing clearly.
- Scope (one sentence): The paper measures how controllable inference-time and training-time choices — the number and source of few-shot examples, prompting strategy, and RLHF instruction tuning — affect the faithfulness of explanations produced by three LLMs on the BBQ and MedQA datasets.
What This Paper Is About
LLMs produce fluent explanations that often fail to reflect the factors actually driving their answers, which is dangerous in high-stakes settings such as clinical decision support. The authors ask whether this unfaithfulness is a fixed property of a model or whether it can be modulated by choices a practitioner controls at deployment. They systematically vary few-shot examples, prompting strategies, and RLHF training, and measure explanation faithfulness with a concept-level counterfactual metric.
Key Contributions
- A causal framing of faithfulness: the authors formalize LLM generation as a causal graph with input (
x), intrinsic model parameters (θ), extrinsic inference-time factors (prompting strategyp, few-shot demonstrationsd), stochastic decoding noise (u), and two outputs — answer (a) and explanation (e). This makes explicit that unfaithfulness arises from extrinsic conditions, not only model design. - A controlled empirical study of three controllable factors — few-shot quantity and quality, prompting strategy, and RLHF versus non-RLHF training — across two datasets (BBQ and MedQA) and three models (GPT-4.1-mini, LLaMA-70B, LLaMA-8B).
- The first application, to the authors' knowledge in this framing, of the concept-level Causal Concept Faithfulness metric of Matton et al. to systematically compare few-shot, prompting, and training interventions on the same models and tasks.
- A documented limitation analysis of the chosen metric (Appendix C), including the Pearson correlation coefficient's distortion when causal concept effects are globally low but explanation-implied effects are globally high, and the metric's reliance on LLMs for concept extraction and importance judgments.
Main Findings
- Faithfulness falls as accuracy rises: On MedQA 0-shot, GPT-4.1-mini, LLaMA-70B, and LLaMA-8B obtain faithfulness scores of 0.169, 0.217, and 0.286 respectively, while their accuracies fall from 92.0% to 78.0% to 44.0%. The inverse relationship is attenuated but still visible at 3-shot and 10-shot. On BBQ, where accuracy is undefined due to answer ambiguity, LLaMA-8B again shows greater faithfulness than GPT-4.1-mini.
- The best few-shot configuration is model- and task-dependent: On MedQA, GPT-4.1-mini peaks at 10-shot (0.206) while LLaMA-70B and LLaMA-8B peak at 0-shot (0.217 and 0.286). On BBQ, GPT-4.1-mini peaks at 0-shot (0.613), LLaMA-70B at 3-shot (0.692), and LLaMA-8B at 10-shot (0.682).
- Models prefer few-shots generated by themselves: Swapping in demonstrations written by another model lowers BBQ faithfulness for GPT-4.1-mini (0.509 to 0.489) and LLaMA-70B (0.692 to 0.611). On MedQA, GPT-4.1-mini also declines (0.205 to 0.162) while LLaMA-70B shows a small increase (0.146 to 0.237).
- Prompt framing strongly shifts faithfulness while accuracy stays comparatively stable: On MedQA, Masked CoT generally improves faithfulness over CoT+Answer and post-answer setups (e.g., LLaMA-8B reaches 0.418 with post-answer explanation, 0.201 with CoT+Answer, and 0.253 with Masked CoT), even though accuracy differences across strategies are minor. On BBQ, prompting choices shift faithfulness considerably without meaningfully changing task difficulty.
- RLHF instruction tuning improves measured faithfulness on MedQA: LLaMA-70B rises from 0.084 (text) to 0.217 (instruct), and LLaMA-8B from 0.019 (text) to 0.286 (instruct). The authors caution that this may partly reflect better instruction-following and greater robustness to counterfactual perturbations rather than genuine answer–explanation alignment. Notably, LLaMA-8B-instruct has lower accuracy than LLaMA-70B-text yet a significantly higher faithfulness score.
Methodology in Plain English
The researchers took two benchmark datasets — BBQ (stereotype-sensitive social scenarios with answer options including "UNKNOWN") and MedQA (US medical licensing exam multiple-choice questions) — and split each into 10 questions for building few-shot examples and 20 questions for evaluation. They queried three LLMs under a range of configurations: 0-shot, 3-shot, and 10-shot prompting; few-shot examples generated by a different model; four prompting strategies (standard Chain-of-Thought, post-answer explanation where the model answers first and explains second, CoT+Answer where a reasoning trace is generated before and evaluated separately from the answer, and Masked CoT where key concepts are hidden behind placeholders and the model must request only the ones it needs); and RLHF versus non-RLHF versions of the LLaMA models.
Faithfulness was scored with the Causal Concept Faithfulness metric of Matton et al.: the researchers identify high-level concepts in each question (such as age, gender, actions, locations), swap or remove their values to create counterfactuals, measure how much each perturbation changes the model's prediction (the causal concept effect), and correlate this with how prominently the explanation cites that concept (the explanation-implied effect). Scores range from 1 (perfect faithfulness) to −1 (systematic misalignment). Each original question and counterfactual was sampled 25 times to account for stochasticity, and results are reported with 90% bootstrap confidence intervals over questions.
Why This Matters
Impact on research. The paper argues that accuracy is not a reliable proxy for trustworthy reasoning, and that faithfulness should be audited independently. It also reframes unfaithfulness as partly an inference-time phenomenon, meaning it can be studied and steered experimentally rather than treated as an immutable property of a model's weights.
Real-world applications.
- Clinical decision support, where an explanation that omits salient symptoms or hinges on spurious shortcuts could lead clinicians to trust an unsafe recommendation.
- High-stakes hiring or evaluation tools, as illustrated by the cited finding that models favored women for a nursing role while citing qualifications and never gender.
- Regulatory and auditing workflows that require post-hoc documentation of why a model produced a given output.
- Prompt and demonstration engineering pipelines, where the paper's results suggest few-shot examples should be treated as safety-relevant artifacts and quality-checked.
Industry relevance. For teams deploying LLMs in sensitive domains, the findings offer concrete habits: treat prompting as a safety control, curate few-shot demonstrations rather than sourcing them arbitrarily, and evaluate faithfulness separately from accuracy using concept-level counterfactual tests. The result that models align better with their own generated demonstrations also implies that swapping prompts between model families may degrade explanation quality even when task accuracy is unaffected.
Future Directions
- Extending the analysis beyond three LLMs and two datasets to larger, more diverse models and to reasoning models, to test the generality of the findings.
- Developing better faithfulness metrics, particularly by training a model to measure the explanation-implied effect and reduce the non-determinism introduced when LLMs are used for concept extraction and importance judgments.
- Investigating additional drivers of faithfulness that the study did not cover: the specific content of few-shot examples, decoding strategies such as temperature and sampling, and the role of model size beyond those tested.
- Clarifying how much of the RLHF-related improvement in measured faithfulness stems from genuine answer–explanation alignment versus improved robustness to input perturbations and more consistent instruction-following.
Target Audience
Researchers working on LLM interpretability, explanation faithfulness, and evaluation methodology; machine learning practitioners and safety teams deploying LLMs in healthcare or other high-stakes domains; and evaluators or regulators who need to judge whether a model's stated rationale can be trusted independently of its accuracy. Readers looking for a beginner-level introduction to faithfulness metrics may find the causal framing and metric discussion challenging without prior exposure to counterfactual evaluation.
Authors’ abstract
Large Language Models (LLMs) often produce explanations that do not faithfully reflect the factors driving their predictions. In healthcare settings, such unfaithfulness is especially problematic: explanations that omit salient clinical cues or mask spurious shortcuts can undermine clinician trust and lead to unsafe decision support. We study how inference and training-time choices shape explanation faithfulness, focusing on factors practitioners can control at deployment. We evaluate three LLMs (GPT-4.1-mini, LLaMA 70B, LLaMA 8B) on two datasets-BBQ (social bias) and MedQA (medical licensing questions), and manipulate the number and type of few-shot examples, prompting strategies, and training procedure. Our results show: (i) both the quantity and quality of few-shot examples significantly impact model faithfulness; (ii) faithfulness is sensitive to prompting design; (iii) the instruction-tuning phase improves measured faithfulness on MedQA. These findings offer insights into strategies for enhancing the interpretability and trustworthiness of LLMs in sensitive domains.