Skip to content
AI.info

jobs

Site Reliability Engineer

About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-nativestartups and global enterprises, from bare metal up through the platform services teams actually buildon. Our culture run

Company
Nscale
Location
Houston; New York; San Francisco; Seattle
Status
Closed
Posted
2026-09-08T17:18:22+00:00

Nscale, which runs a GPU cloud for AI, is hiring an SRE to own automation, tooling and the reliability of production services carrying AI and GPU workloads. The work covers SLOs and SLIs, incident rotation, root-cause analysis and post-incident reviews, plus troubleshooting Linux, networking and distributed systems. The posting asks for three to six years in SRE or systems/software engineering, strong Python or Go, and observability experience; AI/GPU, HPC, InfiniBand, RDMA and Kubernetes are nice to have. Pay is $130,000–$200,000.

Original job posting