Description
NVIDIA is hiring Compute/DL Architecture Performance Optimization Interns to develop high-performance operators for NVIDIA GPU libraries such as cuBLAS, TensorRT, cuDNN, cuSparse, and cuTensor; analyze GPU kernel performance and bottlenecks; design software for kernel authoring and shipping; and apply cutting-edge AI technologies to GPU kernel development workflows. The role requires a computer science or similar degree in progress, strong C/C++ and Python programming, GPU programming and CUDA knowledge, compiler and LLVM/MLIR experience, and strong problem-solving, communication, and teamwork skills.
