jobs
Site Reliability Engineer
ABOUT BASETENBaseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless dev
In inglese
- Company
- Baseten
- Location
- San Francisco; Remote; Toronto; New York; Montreal
- Status
- Open
- Posted
- 2026-05-11T21:33:35.139+00:00
Baseten runs inference infrastructure for AI companies, and this SRE role owns reliability of its multi-cloud Kubernetes platform. The person writes observability as code, defines SLOs and SLIs, leads incident response and post-mortems, turns recurring failures into automated mitigations, and diagnoses latency, memory, GPU utilization and model lifecycle issues. The posting wants deep Kubernetes and infrastructure experience, observability and GitOps tooling, plus runbook and incident-management practice. It says prior ML experience is not required. The role sits on the SRE team and spans engineering and operations.
Original job posting