Vai al contenuto
AI.info

jobs

Software Engineer, Model Inference

About the TeamOur Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them t

In inglese

Company
OpenAI
Location
San Francisco
Status
Open
Posted
2025-02-06T17:40:57.429+00:00

OpenAI's Inference team is hiring a software engineer in San Francisco to make the company's largest models run efficiently in high-volume, low-latency production. The work covers latency, throughput and hardware efficiency of the inference stack, tooling that exposes bottlenecks and instability, and optimizing code and an Azure VM fleet for GPU use. The posting asks for at least five years of software engineering experience, PyTorch and Nvidia GPU familiarity (CUDA, NCCL), HPC technologies such as InfiniBand, MPI or NVLink, and production distributed systems. The role supports research by accelerating inference.

Original job posting