Skip to content
AI.info

jobs

Software Engineer, Inference - Performance Optimization

OpenAI's Inference team is hiring a Software Engineer to analyze and optimize performance across the application, model, and fleet layers of the inference stack. The team uses systems profiling, benchmarking, and quantitative analysis to id

Company
OpenAI
Location
San Francisco, United States
Status
Closed
Posted
2026-09-19T03:33:24.923+00:00

OpenAI's Inference team is hiring a software engineer to analyze and optimize performance across the application, model and fleet layers of its inference stack. The work involves profiling, benchmarking and quantitative analysis to find bottlenecks, cut latency and cost, and forecast capacity for future launches. The engineer will build performance models that turn microbenchmark results into cost-to-serve estimates and improve tooling used by engineering and research teams. The posting asks for first-principles reasoning about distributed systems, model inference and hardware efficiency, plus deep profiling and optimization experience.

Original job posting