Skip to content
AI.info

jobs

Senior Machine Learning Engineer, LLM Inference Optimization

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

Company
Nebius
Location
London, United Kingdom
Status
Open
Posted
2026-09-23T15:14:09+00:00

Nebius is hiring a senior machine learning engineer for its Token Factory inference service, based onsite in London on the Applied AI team. The job is hands-on optimization of LLM and VLM endpoints: deploying and extending engines such as vLLM, SGLang and TensorRT-LLM, building compression workflows (quantization, distillation), adding speculative decoding and KV-cache techniques, and writing benchmarks for latency, throughput and cost per token. Requirements include strong Python and PyTorch, transformer inference internals, and collaboration with kernel and platform engineers.

Original job posting