jobs
Staff Software Engineer, GPU Inference
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv
- Company
- Cerebras
- Location
- Toronto, CAN; Sunnyvale, CA
- Status
- Open
- Posted
- 2026-07-28T15:19:52.049+00:00
Cerebras is hiring a staff software engineer to productionize and optimize a GPU serving stack for large language model inference. The work spans vLLM, AMD ROCm, PyTorch and rack-scale AMD GPU infrastructure: deployment and recovery automation for the accelerator fleet, latency and throughput tuning, cross-layer debugging, and correctness testing. The posting asks for eight or more years of engineering experience, strong C++ and Python, hands-on model-serving work, and distributed-systems debugging. It sits in Cerebras's disaggregated inference systems, pairing AMD GPU prefill with decode on its wafer-scale engine.
Original job posting