Skip to content
AI.info

jobs

Staff Software Engineer, GPU Inference

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv

Company
Cerebras
Location
Toronto, CAN; Sunnyvale, CA
Status
Open
Posted
2026-07-28T15:19:52.049+00:00

Cerebras is hiring a staff software engineer to productionize and optimize a GPU serving stack for large language model inference. The work spans vLLM, AMD ROCm, PyTorch and rack-scale AMD GPU infrastructure: deployment and recovery automation for the accelerator fleet, latency and throughput tuning, cross-layer debugging, and correctness testing. The posting asks for eight or more years of engineering experience, strong C++ and Python, hands-on model-serving work, and distributed-systems debugging. It sits in Cerebras's disaggregated inference systems, pairing AMD GPU prefill with decode on its wafer-scale engine.

Original job posting