Research
Mapping Clinical Doubt: Locating Linguistic Uncertainty in LLMs
Overview Research area: Natural Language Processing / mechanistic interpretability, applied to clinical and medical text. Technical level: Intermediate. The paper assumes familiarity with transformer

- arXiv
- 2511.22402
- Published
- 2025-11-27
- Authors
- Srivarshinee Sridhar, Raghav Kaushik Ravi, Kripabandhu Ghosh
AI summary
Overview
- Research area: Natural Language Processing / mechanistic interpretability, applied to clinical and medical text.
- Technical level: Intermediate. The paper assumes familiarity with transformer layers, residual stream activations, and probing methods, but the central metric is simple and clearly explained.
- Scope: A layer-wise probing study of whether and where three small instruction-tuned LLMs internally represent linguistic (epistemic) uncertainty in clinical statements, using a newly curated contrastive dataset and a proposed metric called Model Sensitivity to Uncertainty (MSU).
What This Paper Is About
Most work on uncertainty in LLMs measures confidence in the model's output — how calibrated its probabilities are, or how truthful it is. This paper asks a different question: does the model's internal activation space change when the same clinical statement is hedged rather than stated confidently? To answer it, the authors build paired prompts that differ only in epistemic modality (for example, "is consistent with" versus "may be consistent with") and measure how far apart the model's internal representations of the two variants sit at each layer.
Key Contributions
- A probing framework for epistemic modality. The authors propose a method for testing whether LLMs encode linguistic uncertainty through measurable shifts in their activation space, rather than through output probabilities alone.
- A new metric: Model Sensitivity to Uncertainty (MSU). MSU is a layer-wise measure defined as the average L2 distance between the activation vectors of a certain and an uncertain variant of each input pair, taken at the final token position of the residual stream. Larger MSU means greater sensitivity to uncertainty at that layer.
- A released contrastive dataset of 3,114 sentence pairs. Sentences were derived from the Anthropic/Persuasion corpus, with modal verbs such as "should" and "must" programmatically masked and replaced by controlled options representing either certain or uncertain modality, yielding 3,114 samples per condition and 6,228 examples in total.
- Evidence of depth-dependent, structurally consistent encoding. Across the three tested models, sensitivity to uncertainty rises with layer depth, suggesting epistemic information is progressively encoded rather than localized in isolated layers.
Main Findings
- Sensitivity increases with depth. In all three models, MSU scores increase monotonically across the transformer stack. Early layers (roughly Layers 0-5) show low sensitivity with MSU around 2.5-6, reflecting surface-level processing, while later layers (roughly Layers 13-23) rise sharply and often surpass 30.
- Concrete layer-wise growth in LLaMA 3.2-1B-Instruct. MSU climbs from 2.67 at Layer 0 to 31.96 at Layer 15. (The Figure 6 caption reports the Layer 0 value as 2.68; the body text states 2.67.)
- Average MSU differs by model. Qwen1.5-0.5B-Chat is the most sensitive of the three with an average MSU of 16.968, followed by Qwen2.5-0.5B-Instruct at 11.520 and LLaMA 3.2-1B at 9.361.
- Qwen variants behave almost identically. Despite differences in architecture and training, Qwen2.5 and Qwen1.5 show nearly identical MSU profiles, which the authors read as evidence that late-emerging sensitivity is robust to model family, scale, and instruction tuning.
- Certain and uncertain inputs are linearly separable. Layer-by-layer PCA of the activation vectors shows clear clustering of certain versus uncertain examples in all three models, validating that the contrastive dataset does probe something real in the representation space.
- A geometric inversion appears in deeper layers. In the later layers of the Instruct model (Layer 17 and 23) and in Layers 13 through 15 of the Chat model, the PCA projection of the uncertain cluster flips position relative to the certain cluster along the primary axes. The authors interpret this reorientation of PC1 and PC2 as a transition from syntactic or lexical representation toward task-relevant abstractions of uncertainty.
- Epistemic uncertainty looks like a high-level semantic feature. The authors connect the late-layer amplification to prior work showing deeper layers handle abstract, compositional semantics and final decision-making.
Methodology in Plain English
The researchers started from claims in the Anthropic/Persuasion corpus. They used pandas, NumPy, and NLTK to find sentences containing modal verbs such as "should" and "must," mask those verbs, and replace them with a controlled multiple-choice option — one certain, one uncertain (for example, "Must" versus "Might"). Appending the letter of the chosen option to the instruction produced a matching pair of prompts, one certain and one uncertain, identical in every other respect. The design deliberately avoids semantically opposite or adversarial pairs; the only systematic difference is the degree of certainty conveyed.
Three small instruction-tuned models were used: Qwen2.5-0.5B-Instruct (391M parameters, 24 layers), Qwen1.5-0.5B-Chat (308M parameters, 24 layers), and Llama-3.2-1B-Instruct (1.1B parameters, 16 layers). Internal activations were extracted with the TransformerLens library, which reads cached internal states without altering the architecture. For each pair, the researchers recorded residual stream activations at the final token position, where the model's output is most strongly influenced, then computed MSU as the average L2 distance between the certain and uncertain activations at each layer. As a diagnostic check on the dataset, they also ran layer-wise PCA with Scikit-learn, projecting certain and uncertain activations onto the two principal components to see whether the two conditions cluster separately.
Why This Matters
Impact on research. The paper shifts attention from output-side uncertainty quantification to input-side representational sensitivity, arguing that calibration metrics alone do not reveal whether a model distinguishes certain from uncertain prompts internally. It also supplies a reusable dataset and a simple, transferable metric for mechanistic interpretability, and it reports a structural pattern that holds across two model families.
Real-world applications:
- Clinical decision support, where a system that treats "may indicate" and "indicates" as equivalent could overstate a diagnosis to a clinician.
- Radiology or pathology report interpretation, where hedging language ("consistent with," "cannot exclude") carries diagnostic weight that downstream systems must preserve.
- Patient-facing health information tools, where responses to hedged questions should appropriately convey uncertainty rather than assert definitive answers.
- Auditing and regulatory review of medical AI, where knowing where a model encodes uncertainty helps reviewers inspect the layers that actually carry it.
Industry relevance. Any organization deploying LLMs in legal, medical, or policy settings — domains the authors explicitly name as high-stakes — needs evidence that models respond to hedged language in a principled way rather than incidentally. Identifying that uncertainty is consolidated in deeper layers gives engineers a concrete target for probing, monitoring, or intervention.
Future Directions
- Broaden the model set. The authors state their findings are based on a limited set of instruction-tuned models and that generality across architectures, sizes, and pretraining paradigms remains uncertain. They call for non-instruct, multilingual, and domain-specific models.
- Broaden the linguistic inputs. Future work should test varied modal verbs, syntax, and discourse structures rather than the masked-modal-verb construction used here.
- Localize uncertainty to components. A key direction the authors name is pinpointing uncertainty to specific neurons or attention heads, moving from layer-level to circuit-level resolution.
- Relate internal signals to output behavior. The paper calls for examining how internal uncertainty signals such as logits and entropy relate to output confidence and calibration.
Target Audience
This paper suits researchers in mechanistic interpretability and NLP who want a lightweight, reproducible probe for how models represent modality; clinical NLP and clinical AI practitioners concerned with how hedging is handled in medical text; and engineers building or auditing high-stakes LLM applications who need to know which layers carry epistemic information. Readers without a background in transformer internals will still follow the core argument, though familiarity with activation spaces and probing will help.
Authors’ abstract
Large Language Models (LLMs) are increasingly used in clinical settings, where sensitivity to linguistic uncertainty can influence diagnostic interpretation and decision-making. Yet little is known about where such epistemic cues are internally represented within these models. Distinct from uncertainty quantification, which measures output confidence, this work examines input-side representational sensitivity to linguistic uncertainty in medical text. We curate a contrastive dataset of clinical statements varying in epistemic modality (e.g., 'is consistent with' vs. 'may be consistent with') and propose Model Sensitivity to Uncertainty (MSU), a layerwise probing metric that quantifies activation-level shifts induced by uncertainty cues. Our results show that LLMs exhibit structured, depth-dependent sensitivity to clinical uncertainty, suggesting that epistemic information is progressively encoded in deeper layers. These findings reveal how linguistic uncertainty is internally represented in LLMs, offering insight into their interpretability and epistemic reliability.