Skip to content
AI.info

Research

RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation

RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation Overview Research area: LLM evaluation reliability and mechanistic interpretability — specifically using internal model s

RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation
arXiv
2510.08931
Published
2025-10-10
Authors
Ashish Kattamuri, Harshwardhan Fartale, Arpita Vats, Rahul Raja, Ishita Prasad

AI summary

RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation

Overview

Research area: LLM evaluation reliability and mechanistic interpretability — specifically using internal model signals (attention patterns, hidden states, confidence trajectories) to tell apart answers produced by memorized recall versus genuine reasoning, as a route to detecting data contamination.

Technical level: Intermediate. The core idea is accessible, but the paper supplies a heavyweight mathematical appendix with per-feature definitions, and several of the "mechanistic" features are explicitly described in the paper as proxies rather than direct computations.

Scope in one sentence: The paper proposes and evaluates RADAR, a framework that extracts 37 features from a language model's internal states and uses an ensemble of four classifiers to decide whether a given prompt was answered by recall or by reasoning, reporting 93% overall accuracy on a 100-example test set.

What This Paper Is About

Data contamination — overlap between an LLM's training data and the benchmarks used to evaluate it — can inflate scores because a model may be reproducing memorized answers rather than solving problems. Existing detection methods typically require access to the training corpus, rely on n-gram overlap or verbatim-output flagging, and break down under paraphrased contamination; crucially, they cannot say whether a model answered a question by recall or by reasoning. RADAR takes a different route: instead of comparing data, it inspects the model's internal computation, on the premise that recall leaves a focused, fast-converging signature while reasoning leaves a distributed, gradually stabilizing one.

Key Contributions

  1. A mechanistic contamination-detection framework. RADAR (Recall vs. Reasoning Detection through Activation Representation) extracts 37 features spanning surface-level confidence and entropy trajectories plus deep mechanistic properties — attention specialization, circuit dynamics, intervention sensitivity, working memory, and causal-effect proxies — from a target model's attention weights and hidden states.

  2. Demonstrated separability of recall from reasoning. The authors report that mechanistic features distinguish the two response types with 93% accuracy, and that a composite Recall Detection Score (RDS) averages 0.933 for recall tasks versus 0.375 for reasoning tasks.

  3. Interpretable signatures of each process. The paper describes recall as characterized by early confidence, faster convergence, focused attention and specialized heads, and reasoning as characterized by broader circuit complexity, distributed attention, and higher activation flow variance.

  4. A detection tool that does not need training-data access. Because it analyzes the model's own computation, RADAR is presented as a complement to external corpus-comparison methods rather than a replacement requiring corpus visibility.

Main Findings

  • Overall accuracy: RADAR achieved 93.0% overall accuracy on the test set, with task-specific performance of 97.7% on recall tasks and 89.3% on reasoning tasks.

  • Perfect on unambiguous and complex cases: 100% (20/20) on clear recall, 100% (20/20) on clear reasoning, and 100% (30/30) on complex reasoning.

  • Weakest on ambiguous prompts: 76.7% (23/30) on challenging/ambiguous cases — the category the paper identifies as most important for detecting contamination, since a reasoning-looking prompt can be solved by memorization.

  • Training behavior: The classifier reached 96.7% cross-validation accuracy during training on a balanced 30-example set (15 recall, 15 reasoning).

  • Feature separation via RDS: Recall tasks showed a mean Recall Detection Score of 0.933 versus 0.375 for reasoning; a scatter plot of RDS against the Reasoning Complexity Index (RCI) reportedly shows clear clustering, with recall in the high-RDS region and reasoning in lower-RDS regions.

  • Surface versus mechanistic contrast: Surface features showed recall prompts producing higher early confidence and faster convergence, while reasoning prompts showed gradual confidence build-up and later stabilization. Mechanistic features showed recall relying on focused attention and specialized heads, while reasoning engaged broader network resources with higher activation flow variance.

  • Top discriminative features: Specialized Heads, Circuit Complexity, and Hidden State Variance are named as the most discriminative features.

  • Target model and feature counts: The target model was microsoft/DialoGPT-medium, configured with output_attentions=True and output_hidden_states=True; the paper counts 37 features total, though the abstract splits them as 17 surface + 20 mechanistic while Appendix D splits them as 16 surface + 21 mechanistic.

Methodology in Plain English

The framework has three parts. First, a Mechanistic Analyzer runs each prompt through a target LLM (DialoGPT-medium) set up to return attention weights and hidden states, then analyzes attention patterns across every head and layer and examines hidden-state dynamics such as variance, norms, and effective rank (computed via SVD).

Second, Feature Extraction produces 37 numbers per prompt. Surface features track how the model's confidence and entropy evolve layer by layer — mean, standard deviation, min, max, range, the layer where confidence peaks, convergence speed, slope, oscillations, early-versus-late confidence, prediction stability, and information-theoretic quantities such as information gain and layer consistency. Mechanistic features capture attention specialization (how many heads fall below an entropy threshold of 1.5), circuit depth and complexity, robustness to component ablation, working-memory proxies, and causal-effect proxies derived from attention entropy.

Third, a Classifier takes the standardized feature vector (z-scored with StandardScaler) and feeds it to an ensemble of four models — Random Forest, Gradient Boosting, SVM, and Logistic Regression. Each outputs a hard label (recall or reasoning) and a probability; the ensemble takes a majority vote and averages probabilities for a final confidence score.

Training used 30 labeled prompts, and evaluation used a separate 100-example test set spanning four categories: clear recall (20), clear reasoning (20), challenging/ambiguous (30), and complex reasoning (30). All features are computable in a single forward pass with no gradients required.

Why This Matters

Impact on research. The paper reframes contamination detection as a question about computation rather than data overlap. If recall and reasoning really do leave distinguishable internal signatures, then evaluation can be audited without holding the training corpus, and benchmark scores can be interpreted with more nuance than a single accuracy number. The work sits at the intersection of mechanistic interpretability and evaluation methodology.

Real-world applications (as implications of the approach):

  • Auditing public leaderboards and benchmark submissions for answers that look memorized rather than solved.
  • Sanity-checking new evaluation datasets before release, to see whether a reasoning-labeled prompt actually elicits reasoning-like internal behavior.
  • Flagging individual evaluation items where a model "knows" rather than "computes" the answer, helping dataset curators find contaminated or leaked items.
  • Adding an internal-signal check alongside existing output-level and corpus-level contamination detection, as part of a layered evaluation stack.

Industry relevance. Model developers and evaluation teams routinely face the question of whether a score is trustworthy. A method that works without access to the training data, produces interpretable per-feature explanations, scales across architectures in principle, and costs only a single forward pass is attractive for routine pre-release model evaluation and for third-party benchmarking.

Future Directions

  • Scaling to larger models. The study used DialoGPT-medium; whether the same signatures hold in much larger models is untested here, and the authors list scaling as future work.
  • Unsupervised detection. All results here come from a supervised classifier trained on 30 labeled examples; unsupervised detection, which would remove the labeling burden, is named as a future direction.
  • Extending to other contamination types. The paper proposes broadening beyond the recall-versus-reasoning distinction.
  • Replacing proxies with true causal measurement. Appendix F states that causal effects are derived from attention entropy rather than actual interventions, that activation patching is approximated via an entropy proxy rather than real patching experiments, that critical components are approximated using specialized head counts, and that working memory is approximated via rank evolution. Converting these proxies into genuine interventional measurements — and improving the 76.7% accuracy on ambiguous cases — are the natural open questions.

Target Audience

Researchers and practitioners in LLM evaluation, benchmarking, and mechanistic interpretability. It is most useful for people who need to judge whether a benchmark score reflects capability or memorization — evaluation engineers at model labs, benchmark maintainers, dataset curators, and academic groups working on contamination detection or interpretability-based auditing. Readers wanting to reproduce or extend the work will need enough ML background to follow the appendix feature definitions and the ensemble classification scheme; the main text alone is readable at an intermediate level. The authors note the code is publicly available via a linked Colab notebook, and that the work does not relate to positions at Meta, LinkedIn, Proofpoint, or the Indian Institute of Science.

Authors’ abstract

Data contamination poses a significant challenge to reliable LLM evaluation, where models may achieve high performance by memorizing training data rather than demonstrating genuine reasoning capabilities. We introduce RADAR (Recall vs. Reasoning Detection through Activation Representation), a novel framework that leverages mechanistic interpretability to detect contamination by distinguishing recall-based from reasoning-based model responses. RADAR extracts 37 features spanning surface-level confidence trajectories and deep mechanistic properties including attention specialization, circuit dynamics, and activation flow patterns. Using an ensemble of classifiers trained on these features, RADAR achieves 93\% accuracy on a diverse evaluation set, with perfect performance on clear cases and 76.7\% accuracy on challenging ambiguous examples. This work demonstrates the potential of mechanistic interpretability for advancing LLM evaluation beyond traditional surface-level metrics.

Read the original paper