jobs
Principal Engineer, AI Inference Reliability
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv
In inglese
- Company
- Cerebras
- Location
- United States and Canada
- Status
- Open
- Posted
- 2025-10-29T02:39:37.801+00:00
Cerebras is hiring a principal-level, hands-on reliability lead for its inference service, the one that serves hosted models on wafer-scale hardware. The role covers SLOs and incident response, fault detection, failover and throttling across regions and data centers, plus internal chaos-testing and fault-injection tooling, and mentoring engineers. The posting asks for a computer science degree, seven or more years in backend, infrastructure or reliability engineering on large distributed systems, and strong Python, C++, Go or Rust. Work is onsite; experience with large-scale AI infrastructure is a bonus.
Original job posting