jobs
ML/Research Engineer, Safeguards
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers,
- Company
- Anthropic
- Location
- San Francisco, CA; New York City, NY
- Status
- Open
- Posted
- 2025-10-09T18:32:16+00:00
Anthropic's Safeguards ML team is hiring ML and research engineers to detect and mitigate misuse of its AI systems. The work covers classifiers for misuse and anomalous behavior, synthetic data pipelines, monitoring for coordinated attacks such as cyber operations and influence campaigns, and agentic safety: threat models, test environments and prompt-injection mitigations. It also includes research on automated red-teaming and adversarial robustness. The posting asks for 4+ years in ML or research engineering, Python and ML systems experience; language modeling, interpretability and reinforcement learning are listed as useful. Pay: $350,000–$500,000.
Original job posting