Vai al contenuto
AI.info

jobs

Senior Site Reliability Engineer — Token Factory (Inference Platform)

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployme

In inglese

Company
Nebius
Location
Amsterdam, Netherlands; Berlin, Germany; London, United Kingdom; Prague, Czech Republic; Remote - Europe
Status
Open
Posted
2025-06-04T06:20:21+00:00

Nebius is hiring a senior site reliability engineer for Token Factory, its foundation-model inference platform inside the Nebius Cloud GPU cloud. The job covers reliability, performance and observability of the whole inference stack: telemetry pipelines for metrics, logs and traces, Kubernetes autoscaling for GPUs, Terraform modules, routing and retry logic, incident response and post-mortems. The posting asks for Kubernetes, Prometheus, Grafana and Terraform fluency, Python or Bash scripting, SLO and alert design, and GPU serving stacks such as vLLM, Triton or Ray. The role is remote across Europe.

Original job posting