jobs
Staff Site Reliability Engineer – Automation and Platform
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv
In inglese
- Company
- Cerebras
- Location
- Sunnyvale, CA; Toronto, CAN
- Status
- Open
- Posted
- 2025-10-03T21:18:58.442+00:00
Cerebras is hiring a staff site reliability engineer to build automation and platform tooling for its AI inference service, which runs on its wafer-scale chips. The role starts with about a month of hands-on operations work, then shifts to declarative GitOps delivery, capacity provisioning and cluster upgrades across datacenters and clouds. It asks for 8+ years in SRE or platform engineering, Argo CD experience, and observability tooling such as Prometheus, Loki, Tempo and Mimir. No 24/7 on-call rotation is required.
Original job posting