Skip to content
AI.info

creators

Jan Leike - AI alignment researcher at Anthropic

Jan Leike is an AI alignment researcher at Anthropic. He co-led OpenAI's Superalignment team until May 2024 and works on scalable oversight.

Jan Leike - AI alignment researcher at Anthropic

Jan Leike is a German AI alignment researcher. He took bachelor's and master's degrees in computer science at the University of Freiburg and a PhD in reinforcement learning theory at the Australian National University under Marcus Hutter, then a postdoctoral fellowship at Oxford's Future of Humanity Institute. At DeepMind he prototyped reinforcement learning from human feedback. He joined OpenAI in 2021 and worked on InstructGPT, ChatGPT and the alignment of GPT-4. In June 2023 he and Ilya Sutskever became co-leads of OpenAI's Superalignment team, formed to align systems smarter than their supervisors. He resigned in May 2024, writing that at OpenAI "safety culture and processes have taken a backseat to shiny products", days before the team was dissolved. Leike joined Anthropic in May 2024 and built its Alignment Science team, working on scalable oversight, weak-to-strong generalisation and jailbreak robustness; his own site still described him as leading it when it was last changed in September 2025. On 8 May 2026 he wrote that he was "starting a new research project at Anthropic". His Google Scholar profile lists Anthropic and is verified at anthropic.com. TIME listed him among the 100 most influential people in AI in 2023 and 2024; he blogs at aligned.substack.com.

Specialization
AI alignment, AI safety, scalable oversight, reinforcement learning from human feedback
Country
United States

Work