Skip to content
AI.info

jobs

Research Scientist, Interpretability

Anthropic’s Interpretability team is working to reverse-engineer how trained language models work, with the goal of developing a mechanistic understanding that can help make advanced AI systems safer. The team studies how neural-network par

Company
Anthropic
Location
San Francisco, United States
Status
Closed
Posted
2026-09-13T03:31:18.739+00:00

Anthropic's Interpretability team reverse-engineers how trained language models work, aiming for a mechanistic understanding that supports AI safety. The research scientist develops methods to recover algorithms stored in model weights, runs experiments from small toy settings up to large models, analyzes features and circuits, builds experiment infrastructure, and publishes findings, working with Alignment Science, Societal Impacts and Pretraining. The posting asks for a research track record, some interpretability experience, Python, and comfort with messy experiments and clear communication. On-site in San Francisco, with remote considered case by case.

Original job posting