Skip to content
AI.info

Research

nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers

Overview Research area: Mechanistic interpretability tooling for transformer language models (software infrastructure rather than a new interpretability technique). Technical level: Intermediate. The

nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
arXiv
2511.14465
Published
2025-11-18
Authors
Clément Dumas

AI summary

Overview

Research area: Mechanistic interpretability tooling for transformer language models (software infrastructure rather than a new interpretability technique).

Technical level: Intermediate. The paper assumes familiarity with transformer internals (attention, MLP blocks, layer norms, unembedding) and with interpretability tooling such as TransformerLens, NNsight, and HuggingFace Transformers.

Scope in one sentence: nnterp is a lightweight wrapper around NNsight that gives researchers a single, standardized interface for reading and writing transformer internals while running the original, unmodified HuggingFace model implementations.

Paper details: Authored by Clément Dumas (MATS; École Normale Supérieure Paris-Saclay, Université Paris-Saclay), arXiv:2511.14465v2 [cs.LG], 14 Dec 2025, licensed CC BY 4.0, presented under the workshop title "Mechanistic Interpretability." Code is at https://github.com/Butanium/nnterp, docs at https://butanium.github.io/nnterp/, install via pip install nnterp.

What This Paper Is About

Interpretability researchers who want to inspect or modify a model's internal activations face a bad choice: use a tool like TransformerLens, which offers a clean consistent interface but requires reimplementing every architecture (risking numerical mismatch with the real model), or use NNsight to run the exact HuggingFace model, which preserves correctness but inherits HuggingFace's inconsistent, per-architecture module naming. The paper's goal is to remove that tradeoff by adding a standardization layer on top of NNsight, so intervention code written once runs identically across many model families without giving up exact HuggingFace behavior.

Key Contributions

  1. A unified API for transformer internals. Accessors such as model.layers_input[layer_idx], model.layers_output[layer_idx], mlps, mlps_input/output, attentions, and attentions_input/output behave the same way across architectures. The paper states this works across 50+ model variants spanning 16 architecture families in the abstract and introduction, while Section 5 states nnterp supports 21 architecture families (see Appendix B for the list).
  2. Automatic validation on model load. Every StandardizedTransformer runs tests verifying (1) module outputs have expected shapes, (2) attention probabilities sum to 1, (3) interventions affect outputs, and (4) layer skip operations preserve causality. The paper reports this caught a HuggingFace transformers 4.54 change — where Qwen and Llama layers began returning activation tensors instead of tuples — on day one.
  3. Built-in implementations of common interpretability methods. Logit lens, patchscope, and activation steering ship with the library and work across all supported models through the same API.
  4. Attention probability access plus a packaged test suite. With enable_attention_probs=True, attention probabilities are readable and writable via model.attention_probabilities[layer_idx], using NNsight's source feature; researchers can verify custom models locally with python -m nnterp run_tests.

Main Findings

  • Standardization is achieved by renaming, not rewriting. nnterp passes a rename dictionary through NNsight's rename argument — for example, GPT-2's transformer.h, attn, and transformer.ln_f become layers, self_attn, and ln_final, whereas LLaMA's structure simply moves model.layers to layers. Modules remain available under their original names, so users can drop down to raw NNsight for advanced work.
  • Tuple-vs-tensor fragility is handled explicitly. The paper notes that model.layers[layer_idx].output may be a tensor or a tuple depending on architecture, while model.layers_output[layer_idx] always gets or sets the output tensor; this is the stated reason to prefer the accessors.
  • Attention probabilities come with caveats. Access requires the slower eager attention implementation, relies on NNsight's source feature to reach an intermediate forward-pass variable, and is described as very implementation sensitive — it may break if a HuggingFace release renames the attention variable, and the hookpoint is not available on all models. Attention probability support is listed as still a work in progress for DbrxForCausalLM, GptOssForCausalLM, Qwen2MoeForCausalLM, and StableLmForCausalLM.
  • Performance is inherited, not measured here. The empirical validation section reports no benchmarks of its own. It states that nnterp adds minimal overhead because it is a thin interface-only wrapper, and cites NNsight's own performance analysis as showing NNsight matches or exceeds TransformerLens speed while using less memory.
  • Prompt tracking is provided as a convenience layer. Prompt.from_strings takes a prompt plus a dictionary of categories and target strings; the tracked tokens for each category are the first tokens of the target strings with and without beginning-of-word, e.g. for ["London", "Lyon"] the tracked tokens could be ["_London", "Lon", "_Lyo", "Ly"]. run_prompts returns per-category probabilities — the illustrative example shows {"target": torch.Tensor([0.7]), "fake": torch.Tensor([0.1])}.
  • Model coverage is documented per class. Appendix B lists the tested classes: BloomForCausalLM, BloomModel, Ernie4_5_MoeForCausalLM, GPT2LMHeadModel, GPTBigCodeForCausalLM, GPTJForCausalLM, Gemma2ForCausalLM, Gemma3ForCausalLM, Gemma3ForConditionalGeneration, GemmaForCausalLM, Glm4ForCausalLM, Glm4MoeForCausalLM, LlamaForCausalLM, MistralForCausalLM, MixtralForCausalLM, OPTForCausalLM, Phi3ForCausalLM, Qwen2ForCausalLM, Qwen3ForCausalLM, Qwen3MoeForCausalLM, SeedOssForCausalLM, SmolLM3ForCausalLM, DbrxForCausalLM, GptOssForCausalLM, Qwen2MoeForCausalLM, StableLmForCausalLM.
  • Custom model support is a first-class path. When automatic renaming fails, researchers define a RenameConfig mapping custom names (e.g. layers_name="custom_layers", attn_name="custom_attention") to the standardized interface, using dot notation for nested modules and lists for alternative names. Attention probabilities for new models require implementing an AttnProbFunction subclass; the appendix walks through a GPT-J example and a four-step process: explore the forward pass with model.scan(), locate where attention probabilities are computed (typically after dropout), implement the hook, and test with dummy inputs.
  • The stated target structure is fixed. All supported models are expected to follow: embed_tokens, layers[i] containing self_attn and mlp, then ln_final, then lm_head.

Methodology in Plain English

The approach is deliberately minimal engineering rather than new science. Instead of re-implementing transformer architectures, nnterp wraps NNsight's LanguageModel class in a new StandardizedTransformer that hands NNsight a dictionary translating each architecture's native module names into one common vocabulary. A configuration system maps each architecture class to its own renaming rules, so users write the same accessor call regardless of the underlying model. On top of the renaming, the library adds I/O accessors that hide whether a module returns a tuple or a single tensor, an optional attention-probability hook built on NNsight's ability to reach intermediate forward-pass variables, and convenience implementations of standard interpretability techniques. To guard against silent breakage, the library runs a validation suite automatically at model load time and ships that same suite so researchers can run it against their own custom models. The authors then document which model classes and which attention implementations pass those checks, and describe how to extend the renaming config and attention hooks to unsupported models.

Why This Matters

Impact on research. The paper argues that separating interface standardization from implementation details makes interpretability research more reproducible: intervention code can be shared and rerun across architectures, enabling the cross-model validation and model-diffing workflows the field is moving toward, while preserving exact HuggingFace behavior instead of relying on a reimplementation. The day-one catch of the transformers 4.54 tuple-to-tensor change illustrates the concrete cost of silent breakage the library is designed to prevent.

Real-world applications.

  • Cross-architecture interpretability studies where the same intervention script must run on GPT-2, LLaMA, Gemma, Qwen, and others without rewriting layer access.
  • Safety and audit work on deployed open-weight models, where auditors need to modify activations (e.g., steering) on the exact released checkpoint rather than on a reimplementation.
  • Internal evaluation of custom or in-house transformers, using RenameConfig and the packaged test suite to confirm the library hooks the right modules.
  • Reproducibility infrastructure for shared research artifacts, so published intervention code runs for other labs on different models.

Industry relevance. Teams that fine-tune or serve many different open-weight architectures benefit from a single inspection API, since it reduces the maintenance burden of parallel per-architecture hook code and flags upstream library changes before they silently corrupt experiments. The paper also positions the tool for scaling to large models where architecture-specific optimizations matter, which is where reimplementations tend to fall short.

Future Directions

  • Automated architecture detection, so new HuggingFace models can be standardized without hand-written rename rules.
  • Support for non-causal and encoder-decoder architectures, which are outside the current layers[i] / self_attn / mlp / lm_head target structure.
  • Broader internal access, specifically attention KQV values, MLP intermediate activations, and MoE router logits.
  • Integration with NNsight itself and support for remote execution via NDIF, so NNsight experiments can run on remote machines.
  • Completing attention probability support for the four listed model classes (DbrxForCausalLM, GptOssForCausalLM, Qwen2MoeForCausalLM, StableLmForCausalLM) and hardening the attention hookpoint against HuggingFace implementation changes.

Target Audience

Interpretability researchers and engineers who already use TransformerLens or NNsight and need to run the same experiments across multiple model families; safety and evaluation teams auditing open-weight models; and research engineers maintaining hook-based tooling for custom or newly released architectures. Readers will get the most out of it with working knowledge of transformer internals and of the Python/PyTorch ecosystem around HuggingFace Transformers.

Authors’ abstract

Mechanistic interpretability research requires reliable tools for analyzing transformer internals across diverse architectures. Current approaches face a fundamental tradeoff: custom implementations like TransformerLens ensure consistent interfaces but require coding a manual adaptation for each architecture, introducing numerical mismatch with the original models, while direct HuggingFace access through NNsight preserves exact behavior but lacks standardization across models. To bridge this gap, we develop nnterp, a lightweight wrapper around NNsight that provides a unified interface for transformer analysis while preserving original HuggingFace implementations. Through automatic module renaming and comprehensive validation testing, nnterp enables researchers to write intervention code once and deploy it across 50+ model variants spanning 16 architecture families. The library includes built-in implementations of common interpretability methods (logit lens, patchscope, activation steering) and provides direct access to attention probabilities for models that support it. By packaging validation tests with the library, researchers can verify compatibility with custom models locally. nnterp bridges the gap between correctness and usability in mechanistic interpretability tooling.

Read the original paper