Description
NVIDIA is hiring an intern to work on one of three specialized GPU/AI software tracks: TensorRT LLM inference optimization using Python and PyTorch, TensorRT compiler graph optimization using C++, or CuTe DSL and CUDA kernel development and optimization using C/C++, CUDA, Python, and related compiler technologies. The role requires pursuing a master’s or doctoral degree in a relevant field and strong problem-solving, curiosity, and passion for GPU computing and deep-learning software performance.
