creators
John Schulman - PPO creator, Thinking Machines Lab
John Schulman co-founded OpenAI, created the PPO algorithm behind RLHF, and is chief scientist at Mira Murati's Thinking Machines Lab.
John Schulman is an American AI researcher who took a physics degree at Caltech in 2010 and a PhD in electrical engineering and computer sciences at UC Berkeley under Pieter Abbeel, completed in 2016. In December 2015, before finishing the doctorate, he co-founded OpenAI with Sam Altman, Elon Musk, Ilya Sutskever, Greg Brockman and others, and went on to lead its reinforcement learning team. He developed Trust Region Policy Optimization and then Proximal Policy Optimization, the 2017 algorithm that became the standard optimiser in the RLHF pipeline used to train ChatGPT; Berkeley's own citation says he "led the reinforcement learning team that developed ChatGPT". He spent nearly nine years at OpenAI and in 2025 received the university's Mark Bingham Award for Excellence in Achievement by Young Alumni. In August 2024 Schulman left OpenAI for Anthropic to work on alignment, then departed after about six months. In February 2025 he announced a new lab with former colleagues: Thinking Machines Lab, founded by Mira Murati, where he is chief scientist. The lab shipped the Tinker fine-tuning API in October 2025 and released the open-weights Inkling models in July 2026.
- Specialization
- reinforcement learning, RLHF, post-training, policy optimization
- Country
- United States