Skip to content
AI.info

jobs

Staff Software Engineer, Inference API

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference serv

Company
Cerebras
Location
Toronto, CAN
Status
Open
Posted
2026-09-23T16:51:14.897+00:00

Cerebras is hiring a staff engineer to build the ML API layer for its disaggregated inference systems, pairing GPU prefill with decode on its Wafer-Scale Engine. The work covers chat completion and streaming APIs, tool calling, structured outputs and multimodal inputs, request routing across backends, performance and correctness validation, and observability. The posting asks for 5+ years of engineering experience, Python and Go plus a systems language, and production API and Kubernetes experience. It is a hybrid role in Toronto.

Original job posting