Skip to main content

AI Computing Software Intern, GPU Kernel Libraries - 2027 at NVIDIA

Department: Intern

Language
Setup
On-site
Location
Huangpu, Shanghai · Beijing
Type
Full-time, Internship
Level
intern
Posted

Description

NVIDIA is hiring Compute/DL Architecture Performance Optimization Interns to develop high-performance operators for NVIDIA GPU libraries such as cuBLAS, TensorRT, cuDNN, cuSparse, and cuTensor; analyze GPU kernel performance and bottlenecks; design software for kernel authoring and shipping; and apply cutting-edge AI technologies to GPU kernel development workflows. The role requires a computer science or similar degree in progress, strong C/C++ and Python programming, GPU programming and CUDA knowledge, compiler and LLVM/MLIR experience, and strong problem-solving, communication, and teamwork skills.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation