Skip to content
AI.info

jobs

Senior Site Reliability Engineer (SRE, Compute Node Team)

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

Company
Nebius
Location
Amsterdam, Netherlands; Remote - Europe
Status
Open
Posted
2026-01-26T15:51:46+00:00

Nebius, an AI cloud provider, is hiring a senior site reliability engineer for its Compute Node team, which builds and runs the cluster scheduler and node-level services behind virtual machines in every cloud region. The work covers debugging Linux user space and kernel space, troubleshooting CPU, memory, NUMA, cgroups and scheduling problems, QEMU/KVM virtualization, building observability, and leading incident response. The posting asks for deep Linux and virtualization experience, container knowledge, and an SRE mindset. The role is remote across Europe.

Original job posting