jobs
Software Engineer, Fleet Infrastructure
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this
In inglese
- Company
- OpenAI
- Location
- San Francisco; New York City
- Status
- Open
- Posted
- 2025-02-13T19:21:22.786+00:00
OpenAI's fleet infrastructure team runs the GPU fleet behind its model training and deployment. This engineer designs, deploys and operates scheduling and quota systems, Kubernetes cluster provisioning and upgrades, CI/CD, and snapshot delivery that shortens model startup times, working with researchers, hardware and business teams. The posting asks for strong programming skills, hyperscale compute experience, Kubernetes and public cloud work (Azure named), and familiarity with AI/ML workloads as a bonus. It is based in San Francisco with three days a week in office.
Original job posting