jobs
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Overview As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutti
- Company
- Nscale
- Location
- Houston; New York; San Francisco; Seattle
- Status
- Open
- Posted
- 2026-09-03T20:22:44+00:00
Nscale is hiring a principal engineer for its AI Infrastructure Operations team, working on-site in one of four listed cities. The role sets acceptance criteria and performance bars for multi-thousand-GPU clusters, runs distributed training and inference jobs as diagnostics, leads root-cause work on cluster failures and performance regressions, and builds automated burn-in and validation pipelines. It requires 10+ years on large-scale compute infrastructure, hands-on AI workload experience with frameworks such as PyTorch or Megatron-LM, networking depth, and production Python.
Original job posting