creators
Paul Christiano: Alignment Research Center Director
Paul Christiano is executive director of the Alignment Research Center, a co-author of the 2017 RLHF paper and a former head of safety at NIST's CAISI.
Paul Christiano is the executive director of the Alignment Research Center, the Berkeley nonprofit he founded, which is developing methods for white-box analysis of neural networks. ARC's own team page also lists him as its president and as a board member. He is a technical advisor at the Center for AI Standards and Innovation inside NIST, and was previously the centre's head of safety, a post he held when it was the U.S. AI Safety Institute; in that job he designed and ran tests of frontier models for capabilities of national security concern. Since 9 September 2026 he has also sat on the OpenAI Foundation board, on its Safety and Security Committee, chaired by Zico Kolter; he says he joined “to support safety oversight”, and OpenAI says he will keep advising the government while recusing himself from OpenAI matters and model evaluations. He took his PhD at UC Berkeley in 2017, advised by Umesh Vazirani — NIST records the degree as computer science, he describes it as statistical learning theory — and holds a BS in mathematics from MIT. Before founding ARC he ran the language model alignment team at OpenAI, where he co-wrote the 2017 paper “Deep reinforcement learning from human preferences” with Jan Leike, Tom Brown and others, the work that put reinforcement learning from human feedback into practice. In 2023 he was named to the expert advisory board of the UK's Frontier AI Taskforce, in its first progress report. TIME named him to its first TIME100 AI list in September 2023, describing him as one of the principal architects of RLHF.
- Specialization
- AI alignment, AI safety evaluations, reinforcement learning from human feedback
- Country
- United States