Vai al contenuto
AI.info

jobs

Senior Applied Scientist, Efficient LLM Inference & Model Optimization

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

In inglese

Company
Nebius
Location
Palo Alto, California, United States
Status
Open
Posted
2026-07-22T22:53:48+00:00

Nebius seeks a Senior Applied Scientist for its Token Factory. The role owns focused research and production optimization projects in efficient LLM and VLM inference. Work includes quantization, QAT, distillation, speculative decoding, KV-cache reuse and compression, long-context inference, MoE routing, and model/runtime co-optimization. The scientist builds prototypes in PyTorch, Triton, or CUDA-adjacent tooling, designs evaluation methodology, partners with MLEs, publishes credible work, and mentors. The posting requires a PhD, a strong publication record, and hands-on Python/PyTorch coding. It is not papers-only: prototypes must become deployed inference components.

Original job posting