Description
NVIDIA is hiring a Site Reliability Engineer to lead technical strategy and build resilient distributed systems, AI agents, and automation for large-scale enterprise products and services. The role covers observability, incident management, systems architecture, networking, Kubernetes, public cloud, and cross-functional collaboration, with responsibilities including mentoring engineers and improving reliability and developer productivity. Candidates need 10+ years of relevant experience, a relevant technical degree or equivalent experience, strong programming and infrastructure-as-code skills, and expertise in observability and cloud platforms. The position is hybrid and offers a base salary of 168,000–270,250 USD for Level 4 and 208,000–333,500 USD for Level 5, plus equity and benefits.
