Vai al contenuto
AI.info

jobs

Senior Site Reliability Engineer -AI Infrastructure Operations

About NscaleNscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-nativestartups and global enterprises, from bare metal up through the platform services teams actually buildon. Our culture runs

In inglese

Company
Nscale
Location
Houston; San Francisco; Seattle
Status
Closed
Posted
2026-04-03T00:27:54+00:00

Nscale is hiring a senior site reliability engineer for the GPU cloud it runs for AI workloads. The job is to own reliability for critical production services, set SLO and incident processes, lead the hardest incidents, and build automation that removes toil, while mentoring other SREs. The posting asks for six to ten years in SRE or systems or software engineering, strong Python or Go, deep Linux, networking and distributed-systems knowledge, hands-on Kubernetes, and prior AI, GPU or HPC operations. On-call rotation; three US locations.

Original job posting