creators
Neel Nanda: Mechanistic Interpretability at DeepMind
Neel Nanda runs Google DeepMind's mechanistic interpretability team and created TransformerLens, the open-source library for reverse-engineering LLMs.
Neel Nanda runs the mechanistic interpretability team at Google DeepMind, whose job, in his words, is to take a trained neural network and reverse engineer the algorithms and structures it has learned. He read pure mathematics at the University of Cambridge, graduating in 2020, interned in quantitative finance at Jane Street and Jump Trading, then spent the year after his degree on AI safety internships at the Future of Humanity Institute, DeepMind and the Centre for Human-Compatible AI. He worked at Anthropic as a language model interpretability researcher under Chris Olah, went independent, and joined DeepMind’s mechanistic interpretability team in 2022. His independent work on grokking, written with Lawrence Chan, Tom Lieberum, Jess Smith and Jacob Steinhardt, was an oral at ICLR 2023, in the conference's top 25%. He created TransformerLens, an open-source library for mechanistic interpretability of GPT-style language models that loads over 15,000 open-weight models; it is credited to him in its README and maintained by others now. He runs a YouTube channel of paper walkthroughs and live research sessions, and mentors researchers twice a year through a MATS stream. In a July 2025 interview he said the most ambitious version of the field, deeply and reliably reading what a model is thinking, has not worked out, and put his own shift as going from “a low chance this is an incredibly big deal” to “a high chance this is a medium big deal”.
- Specialization
- mechanistic interpretability, AI safety, language models
- Country
- United Kingdom