jobs
Site Reliability Engineer - Ops & Automation
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv
- Company
- Cerebras
- Location
- United States and Canada
- Status
- Open
- Posted
- 2025-10-14T20:24:04.05+00:00
Cerebras runs a fast AI inference service on its Wafer-Scale Engine, and this SRE team keeps it running. The role starts with hands-on operations — releases, capacity changes, cluster upgrades — then moves toward building self-service continuous delivery pipelines and internal tooling to cut operational toil. The stack is Kubernetes, Bazel, Prometheus/Grafana/InfluxDB, Python and Go. The posting wants 2–4+ years of SRE with production Kubernetes and observability experience, and notes there is no 24/7 on-call rotation. Locations are the SF Bay Area and Toronto.
Original job posting