jobs
Research Engineer, Frontier Evals & Environments
About the teamThe Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that
In inglese
- Company
- OpenAI
- Location
- San Francisco
- Status
- Open
- Posted
- 2025-04-13T17:25:42.666+00:00
Research engineer on OpenAI's Agent Post-Training team, building evaluation environments and benchmarks that steer the company's largest training runs. Work includes designing reinforcement-learning environments, methods that automatically probe model behavior, measurement research on reliability and variance, and scalable evaluation systems. Public benchmarks from this kind of role include SWE-bench Verified, MLE-bench and PaperBench. The posting asks for grounding in machine learning, software engineering, systems or statistics, plus hands-on experience with LLMs, post-training, evals or agents. San Francisco, on-site.
Original job posting