jobs
Software Engineer, GPU Inference
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv
- Company
- Cerebras
- Location
- United States and Canada
- Status
- Open
- Posted
- 2025-11-25T18:10:40.28+00:00
Cerebras is hiring a software engineer to productionize the GPU prefill path in its disaggregated inference systems, pairing GPU prefill with decode on the Cerebras Wafer-Scale Engine. The work spans vLLM, PyTorch and the AMD ROCm stack on rack-scale AMD GPU hardware: deployment automation, SLOs, fault recovery, benchmarking, and tuning time to first token, throughput, tail latency and KV-cache behavior. It asks for five or more years of engineering experience, strong C++ and Python, and hands-on vLLM-class serving work. The role is on-site in the United States or Canada.
Original job posting