jobs
Research Engineer, Interpretability
Anthropic is seeking a Research Engineer to join its Interpretability team, which works to reverse-engineer how trained language models function and develop a mechanistic understanding of their internal computations. The team’s work support
- Company
- Anthropic
- Location
- San Francisco, United States
- Status
- Closed
- Posted
- 2026-09-19T03:33:53.332+00:00
Anthropic's Interpretability team is hiring a research engineer to build and maintain training and inference infrastructure for mechanistic interpretability work. The job spans instrumented forward and backward passes, activation extraction, steering vectors, profiling, and accelerator-level optimization, plus research-facing tools and safety-audit systems. It asks for substantial software-engineering experience, Python and a second mainstream language, and familiarity with distributed systems, transformer models, PyTorch/CUDA or JAX/XLA, GPU/TPU optimization and large-scale data pipelines. The role is San Francisco-based and hybrid: at least 25% office time, with remote considered case by case.
Original job posting