Skip to content
AI.info

creators

Buck Shlegeris - CEO of Redwood Research

Buck Shlegeris runs Redwood Research, the Berkeley nonprofit behind the AI control agenda and the alignment-faking study done with Anthropic.

Buck Shlegeris is an AI safety researcher who did research and outreach at the Machine Intelligence Research Institute (MIRI) before turning to applied alignment work. In 2021 he co-founded Redwood Research, a Berkeley nonprofit, with Nate Thomas and Bill Zito, serving first as chief technology officer. He is now its chief executive and sits on its board; Ryan Greenblatt is chief scientist. Redwood introduced and has kept pushing "AI control": techniques for deploying AI systems safely even when they may be misaligned and actively trying to subvert the safeguards around them. Its ICML paper "AI Control: Improving Safety Despite Intentional Subversion" set out monitoring protocols for untrusted model agents, and Shlegeris co-authored "Alignment faking in large language models" with Anthropic, which showed a model strategically hiding misaligned intentions during training. Redwood advises AI companies including Google DeepMind and Anthropic and has worked with the UK AI Security Institute.

Specialization
AI control, AI safety, alignment research, model evaluations
Country
United States

Work