Description
NVIDIA is hiring a Principal Site Reliability Engineer to define the technical vision and architecture for reliability across its AI Platform Runtime and related enterprise systems. The role leads distributed-platform architecture, AI agents and intelligent automation, observability, reliability standards, cross-functional reliability programs, incident leadership, and shared platform capabilities, while mentoring senior engineers. Candidates need 15+ years of relevant infrastructure or platform engineering experience, a relevant degree or equivalent experience, and expertise in distributed systems, cloud platforms, programming, infrastructure automation, and observability. The position is hybrid and offers a base salary of 248,000 USD to 396,750 USD plus equity and benefits.
