Skip to content
AI.info

jobs

Senior HPC Engineer, GPU Compute

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

Company
Nebius
Location
Amsterdam, Netherlands; Berlin, Germany; London, United Kingdom; Remote - Europe
Status
Open
Posted
2024-08-27T09:40:48+00:00

Nebius is hiring a senior HPC cluster engineer for its GPU and InfiniBand team, which works on the GPU, InfiniBand and KVM/QEMU layers of its AI cloud. The work covers performance tuning of GPU clusters and InfiniBand networks, troubleshooting GPU and InfiniBand faults, integrating new GPU hardware through Kubernetes, QEMU and KVM, and automating fault detection. The posting asks for 5+ years of system-level software work, 3+ years with Linux, server and PCIe knowledge, and C/C++, Go or Python. Remote within Europe; coding interviews are part of the process.

Original job posting