Description
NVIDIA is hiring an NCX Senior Engineer to support NVIDIA Cloud Partners with advanced Day 2 operations for large-scale NVIDIA accelerated infrastructure. The role focuses on infrastructure health, observability, telemetry, automated detection and remediation, fleet lifecycle administration, operational readiness, and reusable operational frameworks across GPU, CPU, storage, networking, Kubernetes, and AI workloads. Candidates need a relevant advanced degree or equivalent experience, at least 8 years of infrastructure or site reliability engineering experience, and strong expertise in Linux, Kubernetes, distributed systems, observability, automation, networking, and Python or Go.
