jobs
Staff Site Reliability Engineer - AI Platform Runtime
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices
- Company
- NVIDIA
- Location
- US, CA, Santa Clara
- Status
- Open
- Posted
- 2026-09-08T00:00:00+00:00
NVIDIA is hiring a staff-level site reliability engineer for its AI platform runtime. The role leads technical strategy and roadmaps for large cross-functional reliability, scalability and developer-productivity projects, designs distributed systems, and builds AI agents and skills that automate platform operations, alongside observability and automation work. Applicants need 10+ years in SRE, platform engineering or cloud architecture, a computer science or related degree, and strength in Python, TypeScript, JavaScript or Go, infrastructure-as-code tools, OpenTelemetry, Kubernetes and public cloud. Mentoring engineers across teams is part of the job.
Original job posting