Skip to content
AI.info

jobs

Principal Infrastructure Engineer, AI Cluster Performance & Validation

Overview As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutti

Company
Nscale
Location
Houston; New York; San Francisco; Seattle
Status
Open
Posted
2026-09-03T20:22:44+00:00

Nscale is hiring a principal engineer for its AI Infrastructure Operations team, working on-site in one of four listed cities. The role sets acceptance criteria and performance bars for multi-thousand-GPU clusters, runs distributed training and inference jobs as diagnostics, leads root-cause work on cluster failures and performance regressions, and builds automated burn-in and validation pipelines. It requires 10+ years on large-scale compute infrastructure, hands-on AI workload experience with frameworks such as PyTorch or Megatron-LM, networking depth, and production Python.

Original job posting