creators
Evan Hubinger - Anthropic alignment science lead
Evan Hubinger leads alignment science at Anthropic, started its Alignment Stress-Testing team and wrote the Sleeper Agents and alignment-faking papers.
Evan Hubinger leads alignment science at Anthropic. In January 2024 he started the company's Alignment Stress-Testing team, which red-teams Anthropic's own alignment techniques and evaluations by demonstrating empirically how they could fail. He worked before that at MIRI, OpenAI, Google, Yelp and Ripple, and took a BS in mathematics and computer science at Harvey Mudd College in 2019. He is known for the ideas of inner alignment and deceptive alignment, set out in "Risks from Learned Optimization in Advanced Machine Learning Systems" (2019): a trained model can itself become an optimiser pursuing goals its training did not intend. He was lead author of Anthropic's "Sleeper Agents", which built models that behaved helpfully while hiding a backdoored objective and kept it through supervised fine-tuning and reinforcement learning, and co-authored "Alignment Faking in Large Language Models" with Redwood Research, the first empirical case of a model, Claude 3 Opus, strategically complying with a training objective to preserve its existing preferences. In September 2026 he said publicly that he puts the chance of AI killing every human within a decade above 10 per cent.
- Specialization
- AI alignment, deceptive alignment, model organisms, red teaming
- Country
- United States