jobs
Senior Machine Learning Engineer, LLM Inference Optimization
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme
- Company
- Nebius
- Location
- Palo Alto, California, United States
- Status
- Open
- Posted
- 2026-07-22T22:54:04+00:00
Nebius is hiring a senior machine learning engineer for Token Factory, its inference service for frontier models, based onsite in Palo Alto. The role owns optimisation of LLM and VLM endpoints from model artifacts to production: comparing serving engines, deploying and extending vLLM, SGLang or TensorRT-LLM, building quantisation and compression workflows, and benchmarking latency, throughput, GPU memory and cost per token. The posting asks for strong Python and PyTorch, hands-on transformer inference experience, and knowledge of KV-cache, attention and batching bottlenecks. It lists a US base range of $195,200 to $262,200.
Original job posting