jobs
Site Reliability Engineer
About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-nativestartups and global enterprises, from bare metal up through the platform services teams actually buildon. Our culture run
- Company
- Nscale
- Location
- Houston; New York; San Francisco; Seattle
- Status
- Closed
- Posted
- 2026-09-08T17:18:22+00:00
Nscale, which runs a GPU cloud for AI, is hiring an SRE to own automation, tooling and the reliability of production services carrying AI and GPU workloads. The work covers SLOs and SLIs, incident rotation, root-cause analysis and post-incident reviews, plus troubleshooting Linux, networking and distributed systems. The posting asks for three to six years in SRE or systems/software engineering, strong Python or Go, and observability experience; AI/GPU, HPC, InfiniBand, RDMA and Kubernetes are nice to have. Pay is $130,000–$200,000.
Original job posting