jobs
Machine Learning Engineer - Inference
About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large la
In inglese
- Company
- Together AI
- Location
- San Francisco
- Status
- Open
- Posted
- 2024-06-06T19:12:19+00:00
Together AI is hiring a machine learning engineer for its Inference Engine team in San Francisco, on-site and full-time. The job is to build and optimize the production systems and runtime services behind the company's large language model inference, including data ingestion pipelines, code and design reviews, and developer documentation. Requirements include three or more years of production code experience, Python and PyTorch, and high-performance library work, plus low-level systems knowledge of threading, memory, networking and storage. Preferred: familiarity with vLLM, TensorRT-LLM or TGI, speculative decoding, and CUDA/Triton.
Original job posting