jobs
Research Scientist, Interpretability
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers,
In inglese
- Company
- Anthropic
- Location
- San Francisco, CA
- Status
- Open
- Posted
- 2025-11-07T22:47:14+00:00
Anthropic's Interpretability team seeks research scientists to reverse-engineer how trained language models work, focusing on mechanistic interpretability. The work involves designing and running experiments, analyzing features and circuits, building tools to visualize results, and sharing findings internally and publicly. Candidates need a strong scientific research record, some interpretability experience, comfort with open-ended experimental work, and Python. They should collaborate well and treat research and engineering as intertwined. The role is based in San Francisco, with remote work considered case-by-case for exceptional applicants.
Original job posting