jobs
Distributed LLM Inference Engineer
About AnyscaleAt Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of l
- Company
- Anyscale
- Location
- San Francisco; Palo Alto
- Status
- Open
- Posted
- 2026-05-27T23:45:11.194+00:00
Anyscale is hiring a Distributed LLM Inference Engineer to build systems and optimizations for large-scale inference, working with product teams on batch and online inference, integrating Ray Data with an LLM engine, and connecting to vLLM and open-source communities. The posting asks for familiarity with high-throughput, low-latency ML inference, deep learning frameworks such as PyTorch, and distributed systems. Bonus points cover Ray, TensorRT-LLM, deep learning compilers, and GPU or CUDA experience. The role is hybrid across San Francisco and Palo Alto. No salary range is listed.
Original job posting