Skip to content
AI.info

Research

Automatically Finding Rule-Based Neurons in OthelloGPT

Overview Research area: Mechanistic interpretability of transformer language models, using the board game Othello as a controlled testbed. Technical level: Intermediate. Familiarity with neural networ

Automatically Finding Rule-Based Neurons in OthelloGPT
arXiv
2511.00059
Published
2025-10-28
Authors
Aditya Singh, Zihang Wen, Srujananjali Medicherla, Adam Karvonen, Can Rager

AI summary

Overview

Research area: Mechanistic interpretability of transformer language models, using the board game Othello as a controlled testbed.

Technical level: Intermediate. Familiarity with neural networks, linear probes, and basic interpretability concepts helps, but the core idea — training decision trees to describe what individual neurons do — is explained in accessible terms.

Scope in one sentence: This paper presents a decision-tree-based pipeline that automatically discovers MLP neurons in OthelloGPT whose activations follow human-readable logical rules about the game board, and verifies their causal importance through ablation experiments.

What This Paper Is About

Interpretability researchers want to know whether the internal computations of a trained transformer can be described in explicit, human-readable terms. OthelloGPT — a 25M-parameter transformer trained only to predict legal moves in Othello, with no explicit rules given — is a useful testbed because it is complex enough to show rich internal structure yet grounded in rule-based game logic that can be checked against ground truth.

The goal of this paper is to automate the search for "rule-based neurons": MLP neurons whose firing patterns correspond to specific, articulable game conditions, such as a neuron that detects when a diagonal move becomes legal. Prior work identified such neurons by hand; this paper builds a pipeline that finds and describes them automatically, then checks whether those descriptions are causally faithful.

Key Contributions

  1. An automatic method for discovering rule-based neurons in OthelloGPT by training regression and binary decision trees to map board-state features to neuron activations, then converting high-activation decision paths into human-readable logical rules in disjunctive normal form (OR-of-ANDs).

  2. A quantitative comparison against simpler baselines — Lasso (L1) regression and RIPPER — showing that decision trees score highest on R², F1, containment against probe-derived features, and Jaccard similarity.

  3. Causal validation of the discovered rules through two intervention studies: layer-wise replacement of interpretable neurons with their decision-tree surrogates, and fine-grained ablation of pattern-specific neurons across the 60 playable squares.

  4. An open-source Python tool that maps rule-based game behaviors to their implementing neurons, released with code and decision trees on GitHub and an interactive Colab notebook, providing a reproducible benchmark for testing other interpretability methods.

Main Findings

  • Roughly half of layer-5 neurons are rule-describable. The abstract reports R² > 0.7 for 913 of 2,048 neurons in layer 5, described by compact, rule-based decision trees. The remainder are attributed to more distributed or non-rule-based computations.

  • Decision trees beat the baselines. Lasso (L1) regression and RIPPER consistently underperformed decision trees on both R² (regression) and F1 (classification) scores, as well as on the containment and Jaccard metrics computed against probe-identified features (Mine, Empty, Theirs, Flipped, Just Played).

  • Rule-based neurons concentrate in layers 5 and 6. The authors report that "valid-move" neurons in layers 5 and 6 are more interpretable with rules, consistent with prior literature identifying those layers as the valid-move-prediction stage.

  • Replacing neurons with decision-tree surrogates preserves performance. When interpretable neurons (cutoff of 0.7) in whole layers were replaced by their decision-tree outputs, valid-move prediction accuracy remained high and KL divergence of the output logits stayed low. Removing those neurons entirely or substituting mean activations caused a sharp drop in performance.

  • Ablation has a pattern-specific effect. Across all 60 squares, the legal square's probability dropped nearly twice as much in the intervention condition as in the control condition; for the "below 1% of original probability" metric the effect was approximately 4x greater in the intervention condition (0.058 vs. 0.015).

  • Excluding the central squares strengthens the result. Twelve squares adjacent to the middle 2x2 region showed subpar intervention results, suggesting greater shared behavior among them. Averaging only over the 48 squares outside the middle 4x4, the "below 1% of original probability" metric showed roughly a 10x greater effect in the intervention condition (0.077 vs. 0.0076).

  • Full metric table (48 squares, excluding middle 4x4). Logit difference 1.854 (intervention) vs. 0.682 (control); probability difference 0.062 vs. 0.031; clean accuracy 0.99995 vs. 0.99998; corrupted accuracy 0.998 vs. 0.999; accuracy difference 0.0016 vs. 0.00065; below 5% 0.239 vs. 0.040; below 10% 0.344 vs. 0.079.

  • The effect is incomplete. The authors note that the incomplete effect on model behavior suggests other prediction mechanisms remain unaccounted for by their trees.

  • Board-state representations are unique, not redundant. In an appendix experiment, a probe trained on the residual stream at the end of layer 5 over 50k games reached 99.4% accuracy on a 10k-game test set; after ablating the probe directions and retraining, accuracy fell to 33% (random), confirming uniqueness of the board-state representation.

  • Training data scaling saturates around 6,000 games. Binary ground-truth classification trees trained on 60, 600, 6,000, 12,000, and 30,000 games (tested on 500 games) showed performance nearly stabilizing at about 6,000 games, so that size was used for more complex variants.

Methodology in Plain English

The researchers treated each MLP neuron as a target to be explained. For every neuron, they trained a small decision tree whose inputs are concrete facts about the Othello board: whether each of the 64 squares is Mine, Yours, or Empty, which square was played most recently, and which tiles were flipped by that move. The tree learns to predict that neuron's activation value (a regression tree) or simply whether the neuron is "on" or "off," where "on" means the activation exceeds 0.1 of its maximum over the training set (a binary tree). The trees have depth 4, are trained over 6,000 games, and are regularized with a minimum node split count of 100 and a minimum leaf node count of 50.

To turn a tree into a readable rule, the authors collect the decision paths leading to high-activation leaves and treat each path as an AND clause, so the whole neuron becomes an OR of ANDs (a disjunctive normal form). They then simplify the clauses — for example, combining "not E2 is theirs" and "not E2 is empty" into the clearer statement "E2 is mine." This produces descriptions like a neuron that fires when a diagonal becomes legal. Given a query such as "C0 is blank AND D1 is theirs AND E2 is mine," the pipeline automatically surfaces which neurons' rules evaluate to True.

To check the trees are meaningful rather than just predictive, the authors compared them against Lasso regression and RIPPER, measured overlap with features identified by linear probes on board state (containment and Jaccard), and ran causal interventions. In the fine-grained experiment, they took each of the 60 playable squares, defined an intervention pattern and a control pattern (both distinct 3-square patterns that make that square legal), zero-ablated the neurons responding to the intervention pattern, and measured how much the model's probability for that square dropped over a 500-game test set.

Why This Matters

Impact on research. The paper provides a ground-truth benchmark for interpretability methods. Because OthelloGPT's rules are known, researchers can check whether a new method — sparse autoencoders, probes, attribution techniques — recovers the same neuron-rule mappings that decision trees find. It also reframes the question of interpretability from "which features matter" to "what compositional logic a neuron implements," capturing the OR-of-ANDs structure that mirrors game rules.

Potential real-world applications (extrapolated from the paper's framing, not claims it makes):

  • Auditing learned systems in domains with explicit rules, such as compliance or eligibility engines, where you want to know whether a network learned the actual rule or a shortcut.
  • Debugging game-playing and planning agents by locating the specific units responsible for a failure mode.
  • Producing human-reviewable documentation of what parts of a model do, for safety and regulatory review.
  • Serving as a test harness for new interpretability tooling before it is applied to large language models where ground truth is unavailable.

Industry relevance. Teams deploying models in regulated or safety-critical settings often need explanations that map to explicit conditions rather than saliency scores. A pipeline that converts neurons into readable logical clauses, plus a causal check that those clauses matter, is a template for that kind of verification. The released Python tool also lowers the barrier for practitioners and researchers to run their own interpretability methods against a known target.

Future Directions

  • Extending beyond single neurons. The authors note their approach assumes individual neurons are the natural unit of analysis, but OthelloGPT may implement higher-order computations distributed across neuron groups or attention heads. Extending tree-based analysis to multi-neuron subspaces is an explicit suggestion.

  • Handling deeper and time-coupled dependencies. Depth-4 trees may fail on neurons whose activation depends on long conjunctive dependencies (e.g., a length-5 AND clause) or on continuous features, and on features that evolve across timesteps.

  • Accounting for non-rule-based circuits. The authors cite prior work identifying global "flip" circuits that propagate ownership changes across timesteps, which are explicitly not rule-based and fall outside the scope of this analysis, leaving the full picture incomplete.

  • Explaining the central squares. Twelve squares adjacent to the middle 2x2 region showed worse intervention results, suggesting greater shared behavior among those squares — a specific open question about why localization fails there.

Target Audience

Interpretability researchers and mechanistic interpretability practitioners who want a concrete, ground-truth testbed for evaluating their methods; machine learning students looking for a worked example of reverse-engineering a small transformer; and engineers or safety researchers interested in tooling that turns neural network internals into human-readable logical rules with causal verification. Readers with no background in interpretability can follow the high-level argument, but the evaluation metrics (R², F1, containment, Jaccard, KL divergence) assume some machine learning familiarity.

Authors’ abstract

OthelloGPT, a transformer trained to predict valid moves in Othello, provides an ideal testbed for interpretability research. The model is complex enough to exhibit rich computational patterns, yet grounded in rule-based game logic that enables meaningful reverse-engineering. We present an automated approach based on decision trees to identify and interpret MLP neurons that encode rule-based game logic. Our method trains regression decision trees to map board states to neuron activations, then extracts decision paths where neurons are highly active to convert them into human-readable logical forms. These descriptions reveal highly interpretable patterns; for instance, neurons that specifically detect when diagonal moves become legal. Our findings suggest that roughly half of the neurons in layer 5 can be accurately described by compact, rule-based decision trees ($R^2 > 0.7$ for 913 of 2,048 neurons), while the remainder likely participate in more distributed or non-rule-based computations. We verify the causal relevance of patterns identified by our decision trees through targeted interventions. For a specific square, for specific game patterns, we ablate neurons corresponding to those patterns and find an approximately 5-10 fold stronger degradation in the model's ability to predict legal moves along those patterns compared to control patterns. To facilitate future work, we provide a Python tool that maps rule-based game behaviors to their implementing neurons, serving as a resource for researchers to test whether their interpretability methods recover meaningful computational structures.

Read the original paper