jobs
Senior Machine Learning Engineer, LLM Inference Optimization
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme
- Company
- Nebius
- Location
- Zurich, Switzerland
- Status
- Open
- Posted
- 2026-09-24T10:41:29+00:00
Nebius is hiring a senior engineer for its Applied AI team to optimize LLM and VLM inference for the Token Factory, its hosted inference service. The work covers model-compression workflows such as quantization and distillation, serving engines like vLLM, SGLang and TensorRT-LLM, and benchmark harnesses measuring TTFT, TPOT and cost per token. Candidates need strong Python and PyTorch, hands-on deployment or optimization of transformer inference, and quantitative reasoning about latency and throughput trade-offs. The role is on-site in Zurich and pairs closely with kernel and platform engineers.
Original job posting