Skip to content
AI.info

jobs

Training Performance Engineer

About the TeamTraining Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable fron

Company
OpenAI
Location
San Francisco
Status
Open
Posted
2025-10-16T22:51:46.077+00:00

OpenAI's Training Runtime group is hiring a Training Performance Engineer in San Francisco, on a hybrid schedule of three days a week in the office. The work covers profiling end-to-end distributed training runs, finding bottlenecks in compute, communication and storage, raising GPU utilization, writing model graph transforms, and building tooling that tracks MFU, throughput and uptime. Candidates should know Python and C++, have run multi-GPU or HPC training jobs, and have exposure to PyTorch, JAX or TensorFlow; NCCL, MPI, UCX, checkpointing and compiler work help. Relocation assistance is offered.

Original job posting