Vai al contenuto
AI.info

jobs

Principal Engineer, Inference Cloud

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv

In inglese

Company
Cerebras
Location
Sunnyvale, CA
Status
Open
Posted
2025-09-29T12:43:55.551+00:00

Cerebras is hiring a Principal Engineer for its Inference Cloud Platform, the cloud layer behind its inference service. The role owns availability, latency, reliability and multi-region scale, with work on traffic architecture, active-active failover, backpressure, load shedding, SLOs, capacity planning and production code for critical paths. It asks for 10+ years in large-scale distributed systems or cloud infrastructure, strong Go, C++ or Python skills, and experience with high-QPS latency-sensitive systems. ML inference or model serving experience is a plus.

Original job posting