jobs
Software Engineer, Inference - Multi Modal
About the TeamOpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, a
- Company
- OpenAI
- Location
- San Francisco
- Status
- Open
- Posted
- 2025-05-21T18:27:41.612+00:00
OpenAI is hiring a software engineer for its Inference team in San Francisco, on-site, to serve multimodal models — image, audio and other non-text inputs and outputs — in production. The work covers designing and optimizing inference infrastructure for large-scale models, improving throughput and latency, and moving research workflows into reliable services, with attention to GPU utilization, tensor parallelism and hardware abstraction. The posting seeks experience scaling LLM or multimodal inference, GPU-based ML workloads, and inference tooling such as vLLM or TensorRT-LLM. Collaboration with researchers and product engineers is central.
Original job posting