Vai al contenuto
AI.info

jobs

Member of Technical Staff - ML Performance

About Us:AI needs a new infrastructure layer. We're building it at Modal.Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt t

In inglese

Company
Modal
Location
New York; San Francisco
Status
Open
Posted
2026-04-21T00:19:46.977+00:00

Modal is hiring engineers to make machine-learning systems faster at scale. The work covers language and diffusion models, pushing throughput up and latency down, and includes contributing to open-source projects and Modal's container runtime. The posting asks for five or more years of high-performance coding, hands-on work with PyTorch and inference engines such as vLLM or TensorRT, and familiarity with Nvidia GPU architecture and CUDA; applicants are asked to describe a concrete GPU performance improvement they made. Low-level Linux, filesystem and container knowledge is listed as a nice-to-have. The role is on-site in New York or San Francisco.

Original job posting