jobs
Staff Software Engineer, Inference
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and
- Company
- CoreWeave
- Location
- London, England
- Status
- Open
- Posted
- 2026-08-05T19:18:07+00:00
CoreWeave's Inference team runs the Kubernetes-native platform behind large-scale GPU workloads. This staff-level role leads cross-team architecture for request routing, scheduling and GPU resource management, with P99 latency and cost-per-token as targets; the engineer also builds inference optimisations such as speculative decoding and KV-cache reuse and sets benchmarking and observability practice. The posting wants eight or more years on distributed systems, production Kubernetes depth, Go, Python or C++, and hands-on inference experience; open-source work on serving frameworks and CUDA-level GPU tuning are preferred.
Original job posting