jobs
Senior Machine Learning Engineer, LLM Inference Optimization
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme
- Company
- Nebius
- Location
- London, United Kingdom
- Status
- Open
- Posted
- 2026-09-23T15:14:09+00:00
Nebius is hiring a senior machine learning engineer for its Token Factory inference service, based onsite in London on the Applied AI team. The job is hands-on optimization of LLM and VLM endpoints: deploying and extending engines such as vLLM, SGLang and TensorRT-LLM, building compression workflows (quantization, distillation), adding speculative decoding and KV-cache techniques, and writing benchmarks for latency, throughput and cost per token. Requirements include strong Python and PyTorch, transformer inference internals, and collaboration with kernel and platform engineers.
Original job posting