Skip to content
AI.info

jobs

Machine Learning Engineer - Inference

About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large la

Company
Together AI
Location
San Francisco
Status
Open
Posted
2024-06-06T19:12:19+00:00

Together AI is hiring a machine learning engineer for its Inference Engine team in San Francisco, on-site and full-time. The job is to build and optimize the production systems and runtime services behind the company's large language model inference, including data ingestion pipelines, code and design reviews, and developer documentation. Requirements include three or more years of production code experience, Python and PyTorch, and high-performance library work, plus low-level systems knowledge of threading, memory, networking and storage. Preferred: familiarity with vLLM, TensorRT-LLM or TGI, speculative decoding, and CUDA/Triton.

Original job posting