Skip to content
AI.info

jobs

Site Reliability Engineer - Ops & Automation

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv

Company
Cerebras
Location
United States and Canada
Status
Open
Posted
2025-10-14T20:24:04.05+00:00

Cerebras runs a fast AI inference service on its Wafer-Scale Engine, and this SRE team keeps it running. The role starts with hands-on operations — releases, capacity changes, cluster upgrades — then moves toward building self-service continuous delivery pipelines and internal tooling to cut operational toil. The stack is Kubernetes, Bazel, Prometheus/Grafana/InfluxDB, Python and Go. The posting wants 2–4+ years of SRE with production Kubernetes and observability experience, and notes there is no 24/7 on-call rotation. Locations are the SF Bay Area and Toronto.

Original job posting