jobs
Senior Backend Engineer, Inference Platform
About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises,
- Company
- Together AI
- Location
- San Francisco
- Status
- Open
- Posted
- 2025-08-22T18:40:32+00:00
Together AI is hiring a senior backend engineer for its inference platform, which serves generative AI models. The work covers global and local request routing, low-latency load balancing across data centers, auto-scaling, multi-tenant traffic shaping and rate limiting, prefix caching, and profiling for latency and throughput. The posting asks for five or more years on large-scale distributed systems, expert-level Rust, Go, Python or TypeScript, and low-level OS knowledge; Kubernetes and GPU stacks like CUDA, Triton and NCCL are helpful. On-site in San Francisco.
Original job posting