Skip to content
AI.info

jobs

Staff Network Site Reliability Engineer

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

Company
Nebius
Location
United States
Status
Open
Posted
2026-06-29T16:00:19+00:00

Staff Network Site Reliability Engineer at Nebius, where the network underpins its AI cloud platform. The role owns reliability goals for network services and critical paths, improves site readiness, inter-site connectivity and operational standards, leads incident response and postmortems, and builds observability and safer change workflows. Requirements include production Linux, networking fundamentals, high-availability operations, Go or Python automation, and infrastructure tooling such as IaC, CI/CD and containers. Bonus experience includes datapath systems, eBPF/XDP, DPDK, staged rollouts and large-scale telemetry. On-site in the United States, paying $179,500–$224,300 USD.

Original job posting