Description
NVIDIA is hiring a HPC Operations Engineer to troubleshoot support requests, improve deployment automation, configuration management, observability, and operational monitoring, and maintain reliable compute servers in a large-scale HPC environment. The role requires a BS in Computer Science or equivalent experience, at least two years of experience, Linux administration, container and scripting knowledge, cluster configuration management, and strong problem-solving and teamwork skills. Additional Linux, job-scheduler, licensing, Perl, high-speed networking, and distributed-storage knowledge are advantageous.
