jobs
Software Engineer, Inference - Performance Optimization
About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand
In inglese
- Company
- OpenAI
- Location
- San Francisco
- Status
- Open
- Posted
- 2026-04-25T02:24:38.573+00:00
OpenAI is hiring a software engineer for inference performance optimization, on-site in San Francisco. The work spans the inference stack at application, model and fleet layers: profiling, benchmarking and modeling to find latency and throughput bottlenecks, turning microbenchmark results into cost-to-serve estimates, and building tools that let other teams weigh latency, capacity, utilization and cost. The posting asks for first-principles reasoning about distributed systems, model inference and hardware efficiency, and comfort ranging from application behavior to kernels, accelerators, networking and fleet scheduling.
Original job posting