Skip to content
AI.info

jobs

Software Engineer, Inference - Multi Modal

About the TeamOpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, a

Company
OpenAI
Location
San Francisco
Status
Open
Posted
2025-05-21T18:27:41.612+00:00

OpenAI is hiring a software engineer for its Inference team in San Francisco, on-site, to serve multimodal models — image, audio and other non-text inputs and outputs — in production. The work covers designing and optimizing inference infrastructure for large-scale models, improving throughput and latency, and moving research workflows into reliable services, with attention to GPU utilization, tensor parallelism and hardware abstraction. The posting seeks experience scaling LLM or multimodal inference, GPU-based ML workloads, and inference tooling such as vLLM or TensorRT-LLM. Collaboration with researchers and product engineers is central.

Original job posting