Skip to content
AI.info

jobs

Senior Backend Engineer, Inference Platform

About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises,

Company
Together AI
Location
San Francisco
Status
Open
Posted
2025-08-22T18:40:32+00:00

Together AI is hiring a senior backend engineer for its inference platform, which serves generative AI models. The work covers global and local request routing, low-latency load balancing across data centers, auto-scaling, multi-tenant traffic shaping and rate limiting, prefix caching, and profiling for latency and throughput. The posting asks for five or more years on large-scale distributed systems, expert-level Rust, Go, Python or TypeScript, and low-level OS knowledge; Kubernetes and GPU stacks like CUDA, Triton and NCCL are helpful. On-site in San Francisco.

Original job posting