creators
Ethan Perez - Anthropic alignment team lead
Ethan Perez leads the alignment team at Anthropic. He built Constitutional Classifiers, introduced automated red teaming and won an ICML 2024 best paper award.
Ethan Perez leads the alignment team at Anthropic, where he works on reducing existential risks from AI systems. He led the team behind Constitutional Classifiers, a safeguard against universal jailbreaks that, by his account, let Anthropic ship Claude Opus 4 and later models despite their potential to help with weapons development. He also introduced automated red teaming, now used across frontier labs for pre-deployment testing, and worked on retrieval-augmented generation. He co-authored "Debating with More Persuasive LLMs Leads to More Truthful Answers", which won a best paper award at ICML 2024, and has published on sleeper agents, chain-of-thought faithfulness and scalable oversight through debate. He took his PhD at NYU under Kyunghyun Cho and Douwe Kiela, funded by the National Science Foundation and Open Philanthropy, after a bachelor's degree at Rice University, and has spent time at DeepMind, Facebook AI Research, the University of Montreal, Uber and Google. Forbes put him on its 2025 30 Under 30 list in AI.
- Specialization
- AI alignment, red teaming, adversarial robustness, scalable oversight
- Country
- United States