jobs
Researcher, Alignment Interpretability
About the TeamThe Interpretability team studies internal representations of deep learning models. We are interested in using representations to understand model behavior, and in engineering models to have more understandable representations
- Company
- OpenAI
- Location
- San Francisco
- Status
- Open
- Posted
- 2026-09-01T22:01:38.925+00:00
OpenAI's Interpretability team is hiring a researcher to study the internal representations of deep learning models, with an eye to alignment of powerful AI systems. The work covers publishing research on ways to understand network representations, building infrastructure for studying model internals at scale, and steering research toward usefulness or long-term scalability. The posting asks for a Ph.D. or research experience in computer science or machine learning, 2+ years of research engineering, Python proficiency, and background in AI safety or mechanistic interpretability. Work is on-site in San Francisco.
Original job posting