Skip to content
AI.info

Research

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs Overview Research area: Interpretability and representation analysis in large language models — spe

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
arXiv
2601.06599
Published
2026-01-10
Authors
Shivam Adarsh, Maria Maistro, Christina Lioma

AI summary

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs

Overview

Research area: Interpretability and representation analysis in large language models — specifically, how the internal "truth vector" of a statement changes when supporting context is added.

Technical level: Intermediate to Advanced. The paper assumes familiarity with residual stream activations, linear probes, activation steering, and retrieval-augmented generation, and it uses cosine-angle and L2-norm geometry as its core measurement tools.

Scope (one sentence): Across four instruction-tuned LLMs and eight dataset subsets, the paper measures the directional change (theta) and relative magnitude change between truth vectors computed with and without context, layer by layer.

What This Paper Is About

Prior work established that LLMs encode whether a statement is true as a linear direction — a "truth vector" — in their residual stream activations, and that linear classifiers can separate true from false statements in that space. What nobody had characterized is how that geometric structure changes once context is introduced through the prompt. This paper fills that gap by quantifying two things for each statement: the angle (theta) between the truth vector with and without context, and the ratio of the truth vector's L2 norm with context to its norm without context.

Key Contributions

  1. The first characterization of how truth geometry transforms when context is added. The authors state explicitly that their work contributes "the first characterisation of how truth geometry transforms when context is added," addressing a gap left by prior work that tested probe transfer across settings but not geometric change.

  2. Two complementary geometric measures applied layer by layer. They define directional change theta (the angle between truth vectors with and without context, Eq. 5) and relative magnitude (the ratio of squared L2 norms, Eq. 7), and compute both across every layer of every model and dataset.

  3. A three-phase description of directional change. Truth vectors are approximately orthogonal in early layers, converge sharply in early-to-middle layers, and then either stabilize or continue increasing in later layers, with the exact boundaries depending on the dataset.

  4. Evidence that model scale changes the encoding strategy. Larger models (LLaMA-3.1-8B, Mistral-Nemo-12B) distinguish relevant from random context primarily through directional change, while smaller models (Qwen3-4B, SmolLM3-3B) do so through magnitude differences — a pattern the authors connect to representational dimensionality.

Main Findings

  • Three-phase directional pattern (theta). Across all four LLMs, theta remains high (near orthogonal) in early layers, drops sharply in middle layers to a minimum, and then either stabilizes or increases. LLaMA and Mistral begin decreasing around layer 9 and reach minima near layer 15; Qwen and SmolLM show prolonged early phases until layers 14–16 with later minima at layers 20–25.

  • Theta never reaches zero. Even at convergence, models keep distinct representations for statements with and without context.

  • Context generally amplifies true/false separation. Relative magnitude values above 1 mean context increases the separation between true and false representations. In the final layer, LLaMA increases the average relative magnitude across 7 out of 8 datasets; results are mixed for the other models (e.g., Mistral reaches 0.85 on Politifact and 0.87 on ScienceFeedback; SmolLM reaches 0.96 on MF2 and 0.95 on CL-Bill).

  • Middle-layer spikes dominate. Across models, middle-layer relative-magnitude spikes almost always exceed 1, with LLaMA spiking around layers 15–20, Mistral around layers 10–15 (followed by a sharp decline through layers 15–20), Qwen around layer 22 (declining until layer 27), and SmolLM around layers 17–19 with a secondary smaller spike around layers 25–27 that is absent for MF2 and Corporate Lobbying.

  • Conflict produces larger geometric change than alignment. Theta values for ConflictQA-Counter consistently exceed those for ConflictQA-Parametric. The authors suggest that when context aligns with parametric knowledge, both pathways reinforce the same truth direction and theta stabilizes, whereas contradictory context keeps competing signals alive in later layers.

  • Larger models show directional sensitivity to relevance. For LLaMA and Mistral, relevant context generally induces a significantly higher theta than random contexts, particularly on Borderlines, Politifact, ScienceFeedback, MF2 and ConflictQA-Counter. The main exceptions are the Corporate Lobbying datasets from LegalBench.

  • Smaller models show magnitudinal sensitivity to relevance. For Qwen and SmolLM, relative magnitudes are significantly higher for relevant context than random context across most datasets even when theta differences are negative or insignificant. SmolLM shows positive magnitudinal differences in almost all settings despite negative theta differences.

  • A proposed dimensionality explanation. The authors hypothesize that larger models work in higher-dimensional spaces (4096 dimensions for LLaMA-3.1-8B, 5120 for Mistral-Nemo-12B) giving room to represent contexts as distinct directions, while smaller models in compressed spaces (2048 dimensions for SmolLM3-3B, 2560 for Qwen3-4B) face greater directional interference, making magnitude scaling more feasible.

  • Larger representational changes do not equal better utilization. ConflictQA-Counter yields the highest theta values yet involves contradictory information processing, while LegalBench shows minimal differences, which the authors read as models struggling with complex legal text.

  • Geometry does not map cleanly onto output probabilities. The authors checked whether theta and relative magnitude correlate with changes in output probability for "True" and "False" tokens; they find some correlations but not consistently across datasets and models.

Methodology in Plain English

For each statement, the researchers build four prompts: one that supports the statement and one that refutes it, each appearing with and without the relevant context. The model is instructed to complete the generation according to a "Selected Choice" field, and the choice ordering is randomized to remove ordering bias. The ground-truth label in the dataset tells them which completions are true and which are false.

They then extract the residual stream activations used to produce the first output token — the activations at the final prompt position, which aggregate the whole input via causal attention and are unaffected by later generated tokens. The truth vector for a statement is the difference between the true and false activations at that position. They compute this separately for the no-context and with-context conditions.

Two numbers summarize the transformation. The first is theta, the arccosine of the normalized dot product between the with-context and no-context truth vectors — effectively the angle between them, where a large theta means the directions of truth are fundamentally different. The second is the ratio of the squared L2 norm with context to the squared L2 norm without context; values above 1 mean context widens the gap between true and false representations, values below 1 mean it narrows it. Both are averaged over all statements in a dataset for each layer.

To test whether relevance matters, they compare real context against five kinds of random context: random characters, random words sampled from the NLTK English corpus, grammatically valid but incoherent "random salad" sentences, randomly sampled Wikipedia paragraphs, and shuffled contexts from the same dataset. Except for the shuffle condition, random contexts are length-matched to the original context. Statistical significance is assessed with the Wilcoxon signed-rank test at p < 0.05, with Bonferroni-corrected differences reported in the appendix.

Only statements where the model followed instructions across all four prompts were retained. The models used are LLama-3.1-8B-Instruct, Mistral-Nemo-12B-Instruct, Qwen3-4B-Instruct, and SmolLM3-3B, run with greedy decoding via the Huggingface API on NVIDIA A100 and H100 GPUs, requiring approximately 500 GPU hours. The authors also verify causality through interventional experiments where steering along these directions reliably flips model outputs.

Why This Matters

The paper offers a geometric account of what happens inside a model when retrieved or in-context information is supplied — the machinery behind retrieval-augmented generation and in-context learning. Because it focuses on internal representations rather than output behavior, it complements work on probe accuracy and generation accuracy, and it suggests that later-layer contributions to truth representation may be context-dependent rather than uniformly redundant.

Real-world applications:

  • Retrieval-augmented generation: Knowing that relevant context amplifies true/false separation in middle layers, and by how much, informs how retrieved passages are ranked, filtered, or truncated before being placed in the prompt.
  • Fact-checking and claim verification systems: The measured theta separates relevant from random context, offering a potential internal signal for judging whether a passage actually bears on a claim.
  • Knowledge-conflict detection: ConflictQA-Counter produces the strongest and most consistent geometric signals across all four models, which could support flagging when supplied context contradicts a model's parametric knowledge.
  • Domain-specific deployment in law and science: The LegalBench Corporate Lobbying results, where relevant and random context often show no significant difference, warn that models may fail to exploit context in complex legal text — useful for anyone deploying LLMs on regulatory or contractual material.

Industry relevance: The findings apply directly to pipeline design for RAG, to the choice of model scale for context-sensitive tasks (larger models encode relevance directionally, smaller ones magnitudinally), and to interpretability tooling built around contrastive activation vectors, since relative magnitude is reported to be a key hyperparameter for steering strength.

Future Directions

  1. Connect geometry to output behavior. The authors find only inconsistent correlations between theta / relative magnitude and changes in output probability, leaving open what actually links internal geometry to generated text.

  2. Test the dimensionality hypothesis. The claim that larger models use direction while smaller models use magnitude because of representational capacity (4096 and 5120 dimensions versus 2048 and 2560) is framed as a hypothesis, not a tested mechanism, and invites controlled experiments across more scales.

  3. Investigate the LegalBench failure mode. Corporate Lobbying datasets have notably low Flesch readability scores (6.9 and 10.7 compared with 40–50 elsewhere), and the paper's discussion of why relevant and random context behave similarly there is cut off in the provided content.

  4. Extend beyond four models and eight subsets. Whether the three-phase pattern, the conflict-versus-alignment asymmetry, and the direction/magnitude split hold for other families, larger scales, and other context types is not established here.

Target Audience

Interpretability researchers studying how LLMs represent propositional truth and how context interacts with parametric knowledge; RAG and in-context learning engineers who need to know how retrieved passages reshape internal representations; and practitioners deploying LLMs in factual, legal, or fact-checking settings who want a concrete, layer-resolved picture of context sensitivity and its failure modes. Some background in transformer internals and vector geometry is needed to follow the measurement design, though the three headline findings are stated plainly enough for a broader technical readership.

Authors’ abstract

Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also known as truth vectors, have been studied in prior work, however how they change when context is introduced remains unexplored. We study this question by measuring (1) the directional change ($θ$) between the truth vectors with and without context and (2) the relative magnitude of the truth vectors upon adding context. Across four LLMs and four datasets, we find that (1) truth vectors are roughly orthogonal in early layers, converge in middle layers, and may stabilize or continue increasing in later layers; (2) adding context generally increases the truth vector magnitude, i.e., the separation between true and false representations in the activation space is amplified; (3) larger models distinguish relevant from irrelevant context mainly through directional change ($θ$), while smaller models show this distinction through magnitude differences. We also find that context conflicting with parametric knowledge produces larger geometric changes than parametrically aligned context. Collectively, these findings provide a geometric characterization of how context transforms the truth vector in the activation space of LLMs.

Read the original paper