Skip to content
AI.info

jobs

Senior Machine Learning Engineer, LLM Inference Optimization

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

Company
Nebius
Location
Palo Alto, California, United States
Status
Open
Posted
2026-07-22T22:54:04+00:00

Nebius is hiring a senior machine learning engineer for Token Factory, its inference service for frontier models, based onsite in Palo Alto. The role owns optimisation of LLM and VLM endpoints from model artifacts to production: comparing serving engines, deploying and extending vLLM, SGLang or TensorRT-LLM, building quantisation and compression workflows, and benchmarking latency, throughput, GPU memory and cost per token. The posting asks for strong Python and PyTorch, hands-on transformer inference experience, and knowledge of KV-cache, attention and batching bottlenecks. It lists a US base range of $195,200 to $262,200.

Original job posting