jobs
Software Engineer, Model Inference
About the TeamOur Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them t
- Company
- OpenAI
- Location
- San Francisco
- Status
- Open
- Posted
- 2025-02-06T17:40:57.429+00:00
OpenAI's Inference team is hiring a software engineer in San Francisco to make the company's largest models run efficiently in high-volume, low-latency production. The work covers latency, throughput and hardware efficiency of the inference stack, tooling that exposes bottlenecks and instability, and optimizing code and an Azure VM fleet for GPU use. The posting asks for at least five years of software engineering experience, PyTorch and Nvidia GPU familiarity (CUDA, NCCL), HPC technologies such as InfiniBand, MPI or NVLink, and production distributed systems. The role supports research by accelerating inference.
Original job posting