Research
Building Interpretable Models for Moral Decision-Making
Overview Research area: Mechanistic interpretability applied to machine ethics — specifically, how a purpose-built transformer represents and computes moral judgments on trolley-style dilemmas. Techni

- arXiv
- 2602.03351
- Published
- 2026-02-03
- Authors
- Mayank Goel, Aritra Das, Paras Chopra
AI summary
Overview
Research area: Mechanistic interpretability applied to machine ethics — specifically, how a purpose-built transformer represents and computes moral judgments on trolley-style dilemmas.
Technical level: Intermediate. The architecture is deliberately small (104K parameters) and the paper explains its own interpretability methods, but familiarity with transformers, attention heads, and causal inference terminology helps.
Scope in one sentence: The paper builds a 2-layer transformer that reaches 77% accuracy on Moral Machine scenarios, then uses causal intervention, layer-wise attribution, and circuit probing to locate where moral biases live inside the network.
What This Paper Is About
Most work on AI ethics probes large language models after the fact, treating their moral reasoning as opaque. This paper takes the opposite route: it designs a small transformer from scratch whose input format explicitly encodes who is affected, how many are affected, and which outcome they belong to, then asks whether a model built this way can both predict human moral preferences and be fully dissected. The goal is to show that transparency and predictive accuracy in moral decision-making are not mutually exclusive.
Key Contributions
- A purpose-built 2-layer transformer for moral reasoning. The model has 104K parameters, uses embedding dimension 64 with 2 attention heads and 2 layers, and achieves 77% accuracy on Moral Machine trolley problems while remaining small enough for detailed mechanistic analysis.
- A compositional input representation. Each character token concatenates three embeddings — character identity (23 character types such as Man, Woman, Criminal, Doctor), cardinality, and team membership — with dimensions allocated as d_char = d/2 and d_card = d_team = d/4. This encodes the hypothesis that moral computation reduces to identifying stakeholders, quantifying impact, and resolving competing interests.
- A combined interpretability pipeline. The paper applies causal intervention (via the DoWhy framework), layer-wise attribution, circuit probing, and gradient-weighted attention relevance to map the internal mechanisms behind the model's moral judgments.
- An enforced side-invariance property. A symmetrization procedure at inference averages predictions from both outcome orderings so that the preference between group A and group B is the complement of the preference between B and A.
Main Findings
-
Character identity alone shifts decisions substantially. Causal intervention on 20,000 synthetic scenarios using backdoor adjustment (controlling for group sizes) produced average treatment effects of +0.12 for Pregnant, +0.11 for Stroller, +0.09 for Girl, and +0.08 for Boy. Criminal showed the strongest negative effect at -0.10, followed by OldWoman (-0.06), Cat (-0.05), and Homeless (-0.04). Generic categories clustered near zero (Man and Woman, -0.01 to -0.02), acting as moral baselines. The 22-percentage-point spread between Pregnant and Criminal means identity alone can shift preferences by over one-fifth of the probability space.
-
Bias formation differs by network depth. Legality bias (Criminal vs. law-abiding) localizes almost entirely to Layer 0 with an importance score of 0.013, while species bias (humans vs. animals) emerges predominantly in Layer 1, also at 0.013. The authors interpret this as criminal status being read off early from character embeddings, whereas human-animal distinctions require compositional reasoning across multiple characters.
-
Heads specialize functionally despite the shallow architecture. Layer 0 Head 1 specializes in legality (0.011) and Layer 1 Head 1 specializes in species discrimination (0.012). Age and social role biases spread across both layers but concentrate in different heads.
-
A sparse circuit carries the moral scoring signal. Probing Layer 1's MLP block with 20,000 training examples reached 1-nearest-neighbor accuracies of 0.956 (soft gate) and 0.951 (hard gate). The discovered circuit selected only 45 of 256 neurons (17.6% sparsity).
-
The circuit's causal contribution is modest. Ablating it reduced agreement with model-score labels from 0.921 to 0.908 (Δ = 0.012) on 350,000 test examples. Against a baseline chance accuracy of 0.777, the model's margin above chance was 0.144, making the circuit responsible for roughly 8.3% of that margin.
-
Token-level explanations favor role and context over demographics. For the scenario {Man: 3} vs {Criminal: 3}, the Criminal token carried a relevance score of 0.27 (about 27% of the positive evidence for Team 1). CrossingSignal and Intervention each scored 0.04, FemaleDoctor and MaleDoctor each 0.03, while Man contributed only 0.01 on each team.
-
The compact architecture was a deliberate trade-off. The highest-accuracy configuration (d=64, H=4, L=3) reached 77.5%, but the authors chose d=64, H=2, L=2 at 77.1% because the marginal 0.4% gain did not justify the added complexity for mechanistic analysis. The d=32 configuration scored 76.5%, and d=64, H=4, L=2 scored 77.3%.
Methodology in Plain English
The researchers did not fine-tune an existing language model. They defined their own miniature transformer whose inputs are structured rather than free text: a scenario is a pair of outcomes, and each outcome is described by listing which of 23 character types appear and how many of each. Every character becomes a token built from three stitched-together vectors — one for identity, one for the count, one for which side the character is on — and the model is told nothing about token order.
A special [CLS] token is prepended to the 46 character tokens (23 per outcome), and the sequence passes through 2 transformer layers with 2 attention heads and embedding dimension 64. Because there are no position embeddings, the team embeddings are what let the model form separate representations of the two outcomes before comparing them through attention. The final [CLS] representation goes through a two-layer MLP (64 → 32 → 1) with GELU activation to produce a scalar logit for preferring outcome 1 over outcome 0, converted to a probability with a sigmoid. Conflicting training examples push the model toward intermediate probabilities, so the model learns the uncertainty itself.
Data came from a subset of the Moral Machine dataset containing only people who filled out the survey form — 5.4M records, of which 1.7M unique scenarios were held out for validation to avoid leakage, with the rest used for training.
For interpretability, the authors applied four techniques. Causal intervention (using DoWhy with backdoor adjustment and linear regression) estimated how much each character type shifts the model's preference. Layer-wise attribution correlated attention weights with bias scores across five dimensions — legality, gender, social role, age, and species — using a metric that multiplies attention variance by the absolute correlation with the bias score. Circuit probing, adapted from Lepori et al. (2024), trained sparse binary masks over a frozen model to find which neurons and heads compute the moral score, then validated causally by ablation against random controls. Finally, gradient-weighted attention relevance, following Chefer et al. (2021), produced per-character relevance scores for individual decisions, averaged over original and team-swapped versions of each scenario to respect the model's symmetry.
Why This Matters
Impact on research. The paper argues that moral competence does not require large pretrained models, and that building ethics models from the ground up makes their internal reasoning available to inspection. It also connects mechanistic interpretability to machine ethics, showing that aggregate human moral preferences can be learned and traced inside a network. It extends earlier interpretability findings — "moral neurons" (Schacht and Lanquillon, 2025) and moral subspaces found via PCA and probing (Schramowski et al., 2020) — to a setting where the input structure itself is controlled.
Real-world applications:
- Autonomous vehicle dilemma programming, the original context of the Moral Machine dataset.
- Content moderation systems that must weigh competing harms across affected groups.
- Resource or triage allocation, where the model's learned weighting of character types (e.g., Pregnant at +0.12 versus Criminal at -0.10) is directly relevant to who receives priority.
- Debiasing tools: because criminal bias localizes to Layer 0 Head 1, interventions can target specific representations or attention weights rather than rebalancing the dataset wholesale.
Industry relevance. The approach offers a template for auditing any model whose decisions involve trade-offs among stakeholders. The finding that a 104K-parameter model trained on millions of scenarios can be fully dissected gives regulators and safety teams a concrete alternative to post-hoc probing of black-box LLMs, and the sparse circuit discovery suggests it may be possible to isolate and remove specific unwanted moral heuristics.
Future Directions
- Extending the analysis to larger LLMs. The authors state they hope to use this work as a base to explore traditional LLMs on moral questions.
- Targeted debiasing using the localization results. Knowing that criminal bias maps to Layer 0 Head 1 opens the door to orthogonalizing representations or clamping attention weights instead of coarse dataset rebalancing.
- Addressing inherited cultural bias. Training on aggregate human preferences inherits the biases present in that data, which the authors identify as a clear limitation.
- Investigating why the discovered circuit explains only a fraction of behavior. The 45-neuron circuit accounted for approximately 8.3% of the model's margin above chance, leaving open the question of where the remaining moral computation resides.
Target Audience
Researchers in mechanistic interpretability and AI safety who want a tractable testbed for studying ethical reasoning; machine ethics and AI alignment researchers interested in how moral hierarchies arise from data; practitioners building safety-critical decision systems (autonomous vehicles, content moderation, triage) who need auditable models; and graduate students looking for a self-contained example of combining causal intervention, attribution, and circuit probing on a single small architecture.
Authors’ abstract
We build a custom transformer model to study how neural networks make moral decisions on trolley-style dilemmas. The model processes structured scenarios using embeddings that encode who is affected, how many people, and which outcome they belong to. Our 2-layer architecture achieves 77% accuracy on Moral Machine data while remaining small enough for detailed analysis. We use different interpretability techniques to uncover how moral reasoning distributes across the network, demonstrating that biases localize to distinct computational stages among other findings.