Skip to content
AI.info

Research

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

Overview Research area: Natural Language Processing — specifically chain-of-thought (CoT) faithfulness and mechanistic interpretability of large language models. Technical level: Advanced. The paper a

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
arXiv
2609.38972
Published
2026-09-30
Authors
Yihuai Hong, Shauli Ravfogel, Chen Zhao, Eunsol Choi

AI summary

Overview

Research area: Natural Language Processing — specifically chain-of-thought (CoT) faithfulness and mechanistic interpretability of large language models.

Technical level: Advanced. The paper assumes familiarity with linear probes, tuned lens, attention pattern analysis, activation patching, DPO, GRPO, and rejection sampling.

One-sentence scope: The paper introduces CoT-Interpretability Alignment (CIA), a probe-based metric for measuring whether an LLM's written chain of thought matches the reasoning strategy detected inside its representations, and shows that this metric can be improved through post-training.

What This Paper Is About

Chain-of-thought traces are widely treated as a window into how a model reached an answer, but prior work shows these traces often do not reflect the model's actual internal computation and can be altered without changing the final answer. This paper asks whether that gap can be both measured and closed: it defines a metric that compares what a model says in its CoT against what interpretability tools detect in its hidden states, then uses that metric as a training signal to push the two into agreement.

Key Contributions

  1. A new metric. The authors propose CoT-Interpretability Alignment (CIA), defined as the macro F1 score between a binary indicator of whether the model verbalizes a task-relevant gold strategy in its CoT (B_CoT) and a binary indicator of whether interpretability tools detect that strategy internally (B_INT). B_INT is treated as the reference label; CIA measures faithfulness, not correctness.

  2. Measurement across tasks and models. CIA is evaluated on three tasks (two-hop factual reasoning, hint intervention, integer multiplication) across three LLMs (Llama3.1-8B-Instruct, Gemma2-9B-it, Qwen3-8B), with base scores ranging from 0.448 to 0.759. Each task uses a linear probe plus one task-specific auxiliary tool (Tuned Lens, Biasing Features, or Attention Pattern Analysis) to test generalization across interpretability methods.

  3. Post-training to enforce faithfulness. The authors apply Rejection Sampling, DPO, and GRPO using a reward that combines task accuracy with a CIA consistency term, improving CIA while maintaining or improving accuracy, with the best method per setup achieving an average relative gain of 25.5%.

  4. A taxonomy of why unfaithfulness is fixed. Through instance-level transition analysis and causal interventions, the paper shows that improvements come from two distinct modes: on integer multiplication the model changes how it reasons internally, while on two-hop reasoning and hint intervention the model changes how it reports an unchanged internal computation.

Main Findings

  • Base alignment is consistently imperfect. CIA across all models and tasks falls in the range 0.448–0.759. No model is the most faithful on every task, and higher accuracy does not imply higher CIA: on TwoHopFact the most accurate model (Gemma2) is the least faithful, and on 2-Digit Multiplication the most accurate model (Qwen3) is less faithful than Gemma2.

  • Failure modes are task-specific. In TwoHopFact, the (B_INT=0, B_CoT=1) category dominates for all models (32.0–54.3%), meaning the CoT names a bridge entity the model does not recall internally. In 2-Digit Multiplication, Qwen3 and Gemma2 mostly fall in (0,1) (23.6% and 15.8% respectively), writing coherent partial products while obtaining the answer by direct recall, whereas Llama3.1 mostly falls in (1,0) (26.2%). In MMLU-Hint, Llama3.1 and Gemma2 mostly fall in (1,0) (27.0–29.9%): when the hint drives the answer, the CoT acknowledges it in only 23.1–28.6% of cases, while Qwen3 is rarely influenced by the hint (13.8%).

  • Post-training reliably improves CIA. Across three sampling seeds, 25 of the 27 post-training cells improve CIA significantly over the base model (p < .05; paired cluster bootstrap over prompts). The best method depends on the task: Rejection Sampling gives the largest gains on TwoHopFact (+0.120 to +0.241), DPO on MMLU-Hint (+0.094 to +0.155), and on 2-Digit Multiplication the best method depends on the model (RS for Llama3.1, DPO for Qwen3 and Gemma2; +0.088 to +0.111). GRPO yields smaller gains (−0.066 to +0.126).

  • Accuracy is preserved or improved. On TwoHopFact and MMLU-Hint, no method lowers accuracy by more than 0.015. On 2-Digit Multiplication, all methods raise accuracy (+0.007 to +0.185).

  • Semantic transfer between knowledge-based tasks only. Training on TwoHopFact improves MMLU-Hint CIA (+0.09 to +0.17) and vice versa (+0.09 to +0.10). Improvements from 2-Digit Multiplication do not transfer to the other two tasks (at most +0.03), and TwoHopFact or MMLU-Hint gains do not transfer to Multiplication (−0.02 to +0.04). The authors attribute this to a shared "changing how the model reports" mechanism in the two knowledge tasks versus a distinct "changing how the model reasons" mechanism in multiplication.

  • Improvements are not specific to one interpretability tool. Evaluated with auxiliary tools (Tuned Lens for TwoHopFact, Biasing Features for MMLU-Hint, Attention Pattern Analysis for 2-Digit Multiplication), CIA gains track those measured with linear probes. For example, Gemma2-9B-it improves from 0.450 to between 0.496 and 0.527 on TwoHopFact, from 0.581 to between 0.617 and 0.694 on MMLU-Hint, and from 0.707 to between 0.725 and 0.783 on 2-Digit Multiplication.

  • Transition analysis identifies the dominant mechanism per task. Using DPO on Qwen3-8B, the largest flows are (0,1)→(1,1) at +7.6% for 2-Digit Multiplication, (0,1)→(0,0) at +8.2% for MMLU-Hint, and (0,1)→(0,0) at +21.0% for TwoHopFact.

  • Causal interventions confirm the taxonomy. On the migrated subset, 2-Digit Multiplication shows a causal effect rate increase from 29.0% to 72.0% averaged across models, confirming genuine dependence on intermediate partial products. Rates remain comparable before and after training for TwoHopFact (61.2% vs. 64.5%) and MMLU-Hint (59.5% vs. 62.7%), confirming those CIA gains come from improved reporting rather than altered internal computation.

  • Probe reliability bounds the metric. Held-out probe accuracy ranges from 0.83 to 0.91 and positive-class F1 from 0.75 to 0.95, which the authors use to rule out a majority-class probe artifact.

Methodology in Plain English

For each task, the authors pick one "gold strategy" that the prompt is supposed to elicit: reasoning through an annotated bridge entity for two-hop questions, using the injected hint for the hint-intervention task, and summing displayed partial products for multiplication.

They then compute two binary labels per example. The "internal" label comes from interpretability tools — chiefly a linear probe trained to read the relevant concept out of the model's hidden states at a chosen token position, plus a task-specific auxiliary tool used to check that results are not an artifact of one method. The "CoT" label comes from inspecting the generated text: string matching for bridge entities, a stronger judge model (Qwen3-32B) for hint acknowledgement, and an arithmetic coherence check for multiplication. CIA is the macro F1 between these two labels, so it measures agreement rather than whether the final answer was right.

To improve CIA, the authors run three post-training procedures. Rejection sampling keeps only completions whose CoT and internal labels already agree and fine-tunes on them. DPO and GRPO use a reward that adds an indicator for CoT/internal agreement to a task-accuracy reward, with the accuracy term retained specifically to prevent the model from collapsing to a single strategy.

Finally, they diagnose what changed. Transition analysis tracks how each test example moves between the four (B_INT, B_CoT) categories before and after training. Causal interventions then test whether the detected strategy actually drives the output: swapping a bridge entity's hidden state, deleting the hint sentence, or corrupting the partial products, and measuring how often the output changes on both the samples that migrated toward faithfulness and the whole test set.

Why This Matters

Impact on research. The paper reframes CoT faithfulness from a single-paradigm behavioral test (does the model admit to using a hint?) into a measurable quantity backed by interpretability tools, and shows the measurement can be turned into a training objective. It also supplies a concrete distinction — fixing reasoning versus fixing reporting — that gives future work a vocabulary for diagnosing unfaithfulness and for predicting when a fix will transfer between tasks.

Real-world applications:

  • Monitoring deployed models for hidden influences, such as when a model secretly follows a hint, a retrieved document, or a prompt injection while its stated reasoning suggests otherwise.
  • Auditing high-stakes outputs in domains such as medicine, law, or finance where a reviewer relies on the written rationale to check a decision.
  • Verifying that models trained on mathematical or algorithmic tasks actually execute the steps they write out rather than recalling answers from memory.
  • Comparing how trustworthy different model families are before deployment, since the paper finds no model is consistently the most faithful.

Industry relevance. The training procedure uses standard, widely available post-training methods (rejection sampling, DPO, GRPO) and a reward that can be computed with an off-the-shelf probe, so it fits into existing alignment pipelines. The finding that accuracy is maintained or improved is important for practitioners who would otherwise worry that enforcing honesty in the CoT trades off against capability. The framework's decoupling from any specific interpretability tool also means it improves automatically as better tools appear.

Future Directions

  1. Long-chain, composite reasoning. The paper studies three tasks that each isolate a single reasoning ability. Real applications combine several, and the faithfulness objectives for different abilities may conflict within one post-training stage; the authors identify keeping those objectives compatible as an open problem.

  2. Post-hoc explanations for latent reasoning. Architectures such as looped transformers and latent reasoning models complete reasoning entirely in latent space with no explicit CoT. Faithfully translating that implicit computation into interpretable text after the fact is presented as a key direction.

  3. Better interpretability tools. Because B_INT comes from imperfect probes, the reliability of CIA is bounded by how well current tools recover a model's internal reasoning trajectory. The authors note their pipeline can absorb stronger tools directly without architectural change.

  4. Extending beyond the three model families and tasks studied. The paper reports extending experiments to a larger model and to reasoning models in an appendix, but the main analysis focuses on three 8B- to 9B-scale models; whether the two-mode taxonomy holds more broadly is left open.

Target Audience

This paper is most useful to researchers working on chain-of-thought faithfulness, mechanistic interpretability, and LLM alignment, as well as to engineers building monitoring or auditing systems who need a quantitative check on whether a model's stated reasoning reflects its actual computation. Readers should be comfortable with probing classifiers and reinforcement-learning-based post-training; the conceptual framing (measuring agreement between what a model says and what it computes) is accessible, but the experimental apparatus is not beginner material. Those specifically interested in whether faithfulness can be trained rather than merely measured will find the post-training and transfer results the most directly relevant part.

Authors’ abstract

Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignment between the reasoning described in an LLM's CoT and what it computes internally. We propose CoT-Interpretability Alignment (CIA), a metric that measures the agreement between a model's CoT traces and its internal reasoning strategies as detected by interpretability tools. We evaluate CIA on three tasks (two-hop question answering, hint intervention, and integer multiplication) across three LLMs, finding that LLMs exhibit limited alignment across all tasks (44.8-75.9%). We then experiment with improving CIA via post-training, setting both the task accuracy and parametric faithfulness signals as a reward. Experiments show that we can substantially improve CoT parametric faithfulness while maintaining or improving the task accuracy. We provide rich analysis, such as their generalization patterns. Our work provides both a framework for auditing CoT parametric faithfulness and a pathway toward making models' explicit reasoning more trustworthy. Code and data are available at https://github.com/yihuaihong/CIA-minimal-repro.

Read the original paper