jobs
Software Engineer, Model Runtime
About the TeamOpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native si
In inglese
- Company
- OpenAI
- Location
- San Francisco
- Status
- Open
- Posted
- 2026-08-24T16:34:31.338+00:00
OpenAI's Hardware organization is hiring a software engineer to build the model runtime inside an inference engine for frontier models running on the company's custom AI silicon. The work covers scheduling, continuous batching, KV-cache and memory management, distributed execution across chips and racks, and performance tooling, done in close partnership with kernel, compiler and silicon teams. The posting asks for strong systems programming in C++, Rust or Python, understanding of LLM inference, and quantitative reasoning about latency, throughput and utilization. It is hybrid in San Francisco.
Original job posting