Description
NVIDIA is hiring a Senior AI/HPC Engineer for its infrastructure specialist team to deploy, manage, and maintain large-scale AI/HPC systems in Linux-based environments. The role serves as a customer-facing domain expert, handles planning and implementation, provides handover documentation, and collaborates with internal teams on bugs, workarounds, and improvements. Required qualifications include a BS/MS/PhD or equivalent experience in a relevant technical field, at least five years of hardware and software support and deployment experience, Linux system administration, cluster management, scripting, networking, and scheduler expertise; hands-on experience with MPI, NCCL, high-speed networks, automation tools, and Kubernetes is preferred.
