Description
NVIDIA is hiring a Senior Performance Engineer to support performance analysis, telemetry, and optimization for large-scale GPU and CPU clusters used in AI and high-performance computing. The role involves profiling and benchmarking workloads, analyzing networking and collective-communication performance, identifying bottlenecks, developing diagnostic tools, defining test plans, and collaborating with hardware, firmware, networking, systems, and software teams. Candidates need a bachelor's or master's degree in a relevant field, at least five years of experience in performance analysis, systems engineering, or HPC/AI infrastructure, and hands-on expertise with high-performance networking and system performance metrics.
