Skip to content
AI.info

jobs

Staff Site Reliability Engineer - AI Platform Runtime

Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices

Company
NVIDIA
Location
US, CA, Santa Clara
Status
Open
Posted
2026-09-08T00:00:00+00:00

NVIDIA is hiring a staff-level site reliability engineer for its AI platform runtime. The role leads technical strategy and roadmaps for large cross-functional reliability, scalability and developer-productivity projects, designs distributed systems, and builds AI agents and skills that automate platform operations, alongside observability and automation work. Applicants need 10+ years in SRE, platform engineering or cloud architecture, a computer science or related degree, and strength in Python, TypeScript, JavaScript or Go, infrastructure-as-code tools, OpenTelemetry, Kubernetes and public cloud. Mentoring engineers across teams is part of the job.

Original job posting