Skip to content
AI.info

jobs

Senior Applied Scientist, Efficient LLM Inference & Model Optimization

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

Company
Nebius
Location
Palo Alto, California, United States
Status
Open
Posted
2026-07-22T22:53:48+00:00

Nebius seeks a Senior Applied Scientist for its Token Factory. The role owns focused research and production optimization projects in efficient LLM and VLM inference. Work includes quantization, QAT, distillation, speculative decoding, KV-cache reuse and compression, long-context inference, MoE routing, and model/runtime co-optimization. The scientist builds prototypes in PyTorch, Triton, or CUDA-adjacent tooling, designs evaluation methodology, partners with MLEs, publishes credible work, and mentors. The posting requires a PhD, a strong publication record, and hands-on Python/PyTorch coding. It is not papers-only: prototypes must become deployed inference components.

Original job posting