jobs
Principal Engineer, Inference Cloud
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv
In inglese
- Company
- Cerebras
- Location
- Sunnyvale, CA
- Status
- Open
- Posted
- 2025-09-29T12:43:55.551+00:00
Cerebras is hiring a Principal Engineer for its Inference Cloud Platform, the cloud layer behind its inference service. The role owns availability, latency, reliability and multi-region scale, with work on traffic architecture, active-active failover, backpressure, load shedding, SLOs, capacity planning and production code for critical paths. It asks for 10+ years in large-scale distributed systems or cloud infrastructure, strong Go, C++ or Python skills, and experience with high-QPS latency-sensitive systems. ML inference or model serving experience is a plus.
Original job posting