Description
NVIDIA is hiring a Senior Site Reliability Engineer to maintain and improve large-scale Kubernetes clusters supporting DGX Cloud for AI researchers and enterprise clients. The role covers operational reliability, SLO/SLI and error-budget management, service launch and post-launch monitoring, GPU workload optimization across major cloud providers, automation, incident triage, root-cause analysis, and on-call support. Candidates need a BS in Computer Science or related technical field or equivalent experience, at least 10 years of production-service operations experience, and expertise in Kubernetes, containerization, microservices, infrastructure automation, programming, Linux, networking, cloud security, observability, and SRE practices.
