Skip to content
AI.info

jobs

Software Engineer, GPU Inference

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv

Company
Cerebras
Location
United States and Canada
Status
Open
Posted
2025-11-25T18:10:40.28+00:00

Cerebras is hiring a software engineer to productionize the GPU prefill path in its disaggregated inference systems, pairing GPU prefill with decode on the Cerebras Wafer-Scale Engine. The work spans vLLM, PyTorch and the AMD ROCm stack on rack-scale AMD GPU hardware: deployment automation, SLOs, fault recovery, benchmarking, and tuning time to first token, throughput, tail latency and KV-cache behavior. It asks for five or more years of engineering experience, strong C++ and Python, and hands-on vLLM-class serving work. The role is on-site in the United States or Canada.

Original job posting