Vai al contenuto
AI.info

jobs

Software Engineer - Model Performance

ABOUT BASETENBaseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless dev

In inglese

Company
Baseten
Location
San Francisco; Toronto; New York; Montreal; Seattle
Status
Open
Posted
2024-03-28T21:01:24.78+00:00

Baseten's Model Performance team is hiring a software engineer to make large language model inference faster. The work covers implementing and productionizing techniques such as quantization, speculative decoding, KV cache reuse, chunked prefill and LoRA, and debugging performance problems inside TensorRT-LLM, vLLM, sglang, PyTorch and CUDA. Applicants need a degree in computer science or a related field, experience in Python or C++, familiarity with LLM optimization and ML libraries, and a deep understanding of GPU architecture. Hybrid across five North American cities.

Original job posting