jobs
Staff Observability Platform Engineer
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting
- Company
- Nscale
- Location
- US
- Status
- Open
- Posted
- 2026-06-09T13:52:45+00:00
Nscale, a GPU cloud provider for AI, seeks a Staff Observability Platform Engineer in the US. The role builds and evolves observability systems for GPU clusters, AI workloads, and distributed infrastructure. Responsibilities include metrics, logs, traces, alerting, telemetry pipelines, standards, architectural decisions, incident postmortems, and mentoring. The posting asks for 6+ years in SRE, platform, infrastructure, or observability engineering; hands-on Prometheus, Grafana, OpenTelemetry, Kubernetes, Go or Python, and Terraform. Distinctive: observability is treated as a product, with emphasis on GPU, AI/ML, HPC, and large-scale compute environments.
Original job posting