Skip to content
AI.info

jobs

ML/Research Engineer, Safeguards

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers,

Company
Anthropic
Location
San Francisco, CA; New York City, NY
Status
Open
Posted
2025-10-09T18:32:16+00:00

Anthropic's Safeguards ML team is hiring ML and research engineers to detect and mitigate misuse of its AI systems. The work covers classifiers for misuse and anomalous behavior, synthetic data pipelines, monitoring for coordinated attacks such as cyber operations and influence campaigns, and agentic safety: threat models, test environments and prompt-injection mitigations. It also includes research on automated red-teaming and adversarial robustness. The posting asks for 4+ years in ML or research engineering, Python and ML systems experience; language modeling, interpretability and reinforcement learning are listed as useful. Pay: $350,000–$500,000.

Original job posting