jobs
Senior HPC Cluster Engineer
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme
- Company
- Nebius
- Location
- Madrid, Spain; Remote - Europe; United Kingdom
- Status
- Open
- Posted
- 2025-03-14T19:10:34+00:00
Nebius is hiring a Senior HPC Cluster Engineer for its GPU & InfiniBand team, which tunes and hardens the core layer of its AI cloud platform. The job covers GPU cluster and InfiniBand performance tuning, root-cause analysis, adding new GPU hardware through Kubernetes, QEMU and KVM, and automation for monitoring and fault resolution. The posting asks for 5+ years of system-level software development, 3+ years of Linux administration and tuning, server architecture knowledge (PCIe, NICs, kernel) and C/C++, Go or Python. Coding interviews are part of the process.
Original job posting