Description
NVIDIA is hiring a Senior Software Engineer specializing in Deep Learning Inference to design, build, and optimize GPU-accelerated software for large-scale language and generative AI model serving. The role focuses on improving SGLang, vLLM, FlashInfer, and related inference libraries; scaling performance across datacenter GPUs and edge SoCs; and using CUDA, CUTLASS, OAI Triton, NCCL, and other open-source tools. Candidates need a master’s or PhD in a relevant field, at least five years of software development experience, strong C/C++ skills, and experience with DL model inference, profiling, or GPU architecture.
