jobs
Software Engineer - Model Performance
ABOUT BASETENBaseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless dev
- Company
- Baseten
- Location
- San Francisco; Toronto; New York; Montreal; Seattle
- Status
- Open
- Posted
- 2024-03-28T21:01:24.78+00:00
Baseten's Model Performance team is hiring a software engineer to make large language model inference faster. The work covers implementing and productionizing techniques such as quantization, speculative decoding, KV cache reuse, chunked prefill and LoRA, and debugging performance problems inside TensorRT-LLM, vLLM, sglang, PyTorch and CUDA. Applicants need a degree in computer science or a related field, experience in Python or C++, familiarity with LLM optimization and ML libraries, and a deep understanding of GPU architecture. Hybrid across five North American cities.
Original job posting