Vai al contenuto
AI.info

jobs

Senior Machine Learning Engineer, LLM Inference Optimization

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

In inglese

Company
Nebius
Location
Zurich, Switzerland
Status
Open
Posted
2026-09-24T10:41:29+00:00

Nebius is hiring a senior engineer for its Applied AI team to optimize LLM and VLM inference for the Token Factory, its hosted inference service. The work covers model-compression workflows such as quantization and distillation, serving engines like vLLM, SGLang and TensorRT-LLM, and benchmark harnesses measuring TTFT, TPOT and cost per token. Candidates need strong Python and PyTorch, hands-on deployment or optimization of transformer inference, and quantitative reasoning about latency and throughput trade-offs. The role is on-site in Zurich and pairs closely with kernel and platform engineers.

Original job posting