Description
NVIDIA is hiring a Senior HPC DevOps Engineer to design, implement, and maintain large-scale HPC and AI clusters, including monitoring, logging, alerting, infrastructure as code, CI/CD, automation, networking, and troubleshooting. The role requires a bachelor's degree in a relevant field, at least five years of experience, and expertise in programming, Kubernetes, container technologies, event streaming, storage systems, virtualization, and cloud platforms, with knowledge of GPU architecture and job scheduling tools such as Slurm and Kubernetes.
