Research
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
Overview Research area: Meta-science and AI-assisted research evaluation, instantiated in mechanistic interpretability; submitted under cs.CY (AI Safety & Ethics). Technical level: Intermediate. The f

- arXiv
- 2602.18458
- Published
- 2026-02-05
- Authors
- Xiaoyan Bai, Alexander Baumgartner, Haojia Sun, Ari Holtzman, Chenhao Tan
AI summary
Overview
Research area: Meta-science and AI-assisted research evaluation, instantiated in mechanistic interpretability; submitted under cs.CY (AI Safety & Ethics).
Technical level: Intermediate. The framework and results are conceptual, but the paper assumes familiarity with mechanistic interpretability work (circuits, activation patching, probe analysis) and with standard ML reproducibility practices.
Scope: The paper proposes an execution-grounded evaluation framework and an agent, MechEvalAgent, that judges research outputs by running their code and data rather than reading only the paper, and validates it against human experts on 30 mechanistic interpretability research outputs.
What This Paper Is About
Peer review treats the paper's narrative as the object of evaluation, but empirical AI research cannot be fully assessed without executing its code and data. This gap is widening because agentic research systems now generate large volumes of outputs autonomously, and because the number of paper submissions to venues like NeurIPS rose from 1,420 in 2013 to 21,575 in 2025. The authors build an agent that verifies the science itself, not just the story told about it, by checking coherence, reproducibility, and generalizability of research outputs that bundle narrative with executable artifacts.
Key Contributions
-
First execution-grounded evaluation framework for research outputs. The framework standardizes research outputs into two components, narrative (plan and report, plus human prompts and research trace for agent-generated work) and execution resources (implementation with code and data, plus a walkthrough), enabling systematic assessment beyond narrative-alone review.
-
MechEvalAgent, an automated evaluation agent built on Claude Code with execution and logging support from Scribe (Goodfire AI, 2025), implementing the framework through three components: coherence (consistency and instruction following), reproducibility (execution quality and replication quality), and generalizability.
-
Empirical evaluation on 30 mechanistic interpretability research outputs, split into three settings of ten examples each: replication of research, open-ended research questions, and human-written repositories.
-
Demonstration that execution matters: MechEvalAgent achieves above 80% agreement with human experts, captures 67 of 87 failures humans identified, and surfaces 51 additional issues humans missed, while ablations without code access or execution drop to about 45% agreement on average.
Main Findings
-
Failures are pervasive. Over 90% of tasks have at least one failure in reproducibility, largely driven by execution errors, and 80% of tasks fail in coherence, mainly due to lack of consistency. Execution failures stem from model loading, shape calculations, and environment setup, and appear in both agent-generated and human-written repositories.
-
Agent evaluations are rated high quality. Human experts scored the quality of MechEvalAgent's evaluation outputs above 4.7 out of 5 across all dimensions on a 1-5 Likert scale.
-
Agreement exceeds 80% across all dimensions, and most human-identified failures are captured (67 of 87).
-
51 additional issues are surfaced that humans miss, concentrated in reproducibility and generalizability, the dimensions that are more time-consuming for humans to evaluate. MechEvalAgent surfaces more unique issues across all dimensions: 22.6% in coherence, 55.3% in reproducibility, and 57.1% in generalizability.
-
Execution quality checks fail most often. Consistency checks fail 27.3% on average, while execution quality checks fail 60% of the time. Plan-implementation consistency (CS2) fails at 23.3%, effect size (CS3) at 6.7%, and lack of sufficient justification at 23.3%.
-
Concrete reproducibility failures. In the acronyms replication task, the research agent reported using the logit correlation metric on the full dataset but actually evaluated only the first 20 examples; MechEvalAgent found the metric deviates by more than 8% from the original (original value 0.66, replication value 0.72) across all three evaluation runs. In the erasing repository (Gandikota et al., 2024), probe-training code and adversarial testing scripts are missing. In the belief repository (Prakash et al., 2025), a Jupyter notebook related to BigToM causal model experiments is invalid and not runnable.
-
Concerning coherence failures. On the ioi task, the report concluded a circuit "strongly supports" the hypothesis despite a -4.2% circuit performance. On the unanswerable task, MechEvalAgent flagged that the final circuit size lacked justification and that circuit ablation was only 0.73× as impactful as random ablation, contradicting the claimed importance.
-
Ablations confirm the value of execution. The Doc-Only evaluator (final report only) and the No-Execution evaluator (full repository but no execution) agree with humans far less, especially on execution-dependent dimensions, and tend to over-assign FAIL.
-
Efficiency. Human judges spent an average of 2.2 hours per evaluation task, with human-written repositories often exceeding 3.35 hours. MechEvalAgent typically completes evaluations in under 30 minutes for agent-generated tasks and around one hour for human-written repositories.
-
Humans still add value. 20 issues were identified exclusively by human experts. In the multilingual task, humans caught a hallucination where a reported support value of +0.13 for a neuron came only from a single translation task rather than the average across the full task. In the uncertainty task, humans criticized an unjustified switch to GPT-2 Medium, but code inspection showed an
ifcondition that switches models only after a load failure, meaning MechEvalAgent's assessment was more accurate. -
Disagreements cluster in specific checklist items. Reproducibility disagreements often come from the Redundancy check (C3), where code blocks copied repeatedly across inputs are marked redundant by evaluators but not by humans. Generalizability disagreements arise for method generalizability (GT3, what counts as a new method) and model generalizability (GT1, evaluation depth).
Methodology in Plain English
The authors first define a unified format for research outputs: a narrative part (plan and report; for human papers, the plan is extracted from the paper) and an execution part (code, data, and a walkthrough). All 30 outputs are standardized into this format.
They then build MechEvalAgent around three evaluation components, each handled by a dedicated agent that works through a structured binary checklist and also produces a longer analysis report:
- Coherence checks consistency (does the implementation match the stated plan, do the results support the claims, are effect size, justification, and statistical significance adequate) and instruction following (did the agent pursue the specified hypotheses and constraints).
- Reproducibility checks execution quality (does the code run and implement the described computations) and replication quality. For replication, a fresh session is run without access to the report, to prevent copying conclusions or reverse-engineering from reported findings. A separate verification agent checks result fidelity and looks for external references or memorized information.
- Generalizability checks whether findings transfer to a new model, new data instances, or related tasks, allowing up to three trials to find suitable alternatives.
To validate, three human authors acted as annotators using the exact same checklists and instructions, including reproducing repositories and testing generalizability. Each of the 30 generations was evaluated by exactly one expert, and human evaluation was conducted on the first evaluation run rather than all three. Each automated evaluation was run three times, with AND logic for PASS (a task passes only if all runs pass). Human experts also rated the quality of the agent's evaluation output from 1 to 5.
Two ablations isolate the value of execution: a Doc-Only evaluator that reads only the final report and cannot inspect or run code, and a No-Execution evaluator that sees the full repository but cannot execute it. Safeguards include strict file access restrictions, GitHub version control monitoring to detect unintended source modifications, and routing replication results through a separate verification agent.
Why This Matters
Impact on research: The paper argues for a division of labor between human and automated review. Human reviewers excel at judging novelty, contribution, and framing, areas where AI is 10 times less likely to comment (Liang et al., 2024); the 20 issues identified exclusively by humans involved interpretation and scope judgments. Automated evaluation excels at execution-heavy checks that humans skip under deadline pressure. Passing MechEvalAgent does not guarantee high quality, but failing it highlights concrete, actionable weaknesses, and its rationales are more specific than the general rationales humans often give, such as "insufficiently justified." The framework is described as not domain-specific and adaptable to other areas of scientific research.
Real-world applications:
- Pre-submission or pre-external-review screening of ML papers and repositories, catching basic execution errors before reviewers see them (aligned with BaHammam, 2025).
- Auditing outputs of autonomous research agents, detecting hallucinations where positive results are reported despite weak evidence.
- Verifying replication claims by running code in a fresh session with report access withheld.
- Supporting journal and conference editors who lack the time or resources to run shared code.
Industry relevance: As agentic workflows are deployed for code assistance, ideation, and end-to-end research, they generate more submissions into an already strained system. The paper notes authors have even inserted prompts into papers to raise AI judge scores. Tooling that grounds evaluation in execution is directly relevant to labs and venues building automated or assisted review pipelines.
Future Directions
- Model ensembles for evaluation. The current implementation uses Claude Code as the sole evaluation agent; the authors suggest exploring ensembles of different models to improve robustness.
- More independent human review. Due to resource constraints, each repository was reviewed by a single human expert, and human evaluation covered only the first of three runs. Additional reviewers would further validate the observed patterns.
- Graded rather than binary judgments. Binary checklists are brittle when evidence is mixed; graded judgments could better reflect uncertainty, particularly because research is highly open-ended and the underspecified space may be larger than in other domains.
- Better instruction following and uncertainty handling. The evaluation agent sometimes modified source files despite prohibitions, downplayed negative evidence in ambiguous settings, or marked error-fixing code as redundant even when its own rationale acknowledged the fix. The authors call for continued work on these behaviors and on clearer guidance for handling uncertainty.
- Extending beyond mechanistic interpretability to other scientific domains, which the framework's modular design is intended to permit.
Target Audience
Peer reviewers, area chairs, and conference or journal organizers dealing with submission growth and reproducibility problems; researchers in mechanistic interpretability and AI safety interested in how agent-generated work is evaluated; developers of autonomous research agents who want a feedback signal for improving reliability; and meta-science researchers studying reproducibility and evaluation methodology. The paper is also useful for readers who want a concrete example of an evaluation agent that inspects artifacts rather than text, since it documents failure modes, disagreement patterns, and safeguards in detail.
Authors’ abstract
Reproducibility crises across sciences highlight the limitations of the paper-centric review system in assessing the rigor and reproducibility of research. AI agents that autonomously design and generate large volumes of research outputs exacerbate these challenges. In this work, we address the growing challenges of scalability and rigor by flipping the dynamic and developing AI agents as research evaluators. We propose the first execution-grounded evaluation framework that verifies research beyond narrative review by examining code and data alongside the paper. We use mechanistic interpretability research as a testbed, build standardized research output, and develop MechEvalAgent, an automated evaluation framework that assesses the coherence of the experimental process, the reproducibility of results, and the generalizability of findings. We show that our framework achieves above 80% agreement with human judges, identifies substantial methodological problems, and surfaces 51 additional issues that human reviewers miss. Our work demonstrates the potential of AI agents to transform research evaluation and pave the way for rigorous scientific practices.