Description
NVIDIA is hiring a Senior Software Engineer specializing in Deep Learning Inference to design, build, and optimize GPU-accelerated software for large-scale model serving and inference. The role focuses on improving open-source inference frameworks, optimizing LLM and generative AI models across datacenter GPUs and edge SoCs, and contributing to libraries such as vLLM, SGLang, FlashInfer, and LLM software solutions. Candidates need a master’s or PhD or equivalent experience, at least five years of relevant software development experience, strong C/C++ skills, and familiarity with CUDA, Triton, CUTLASS, or related GPU technologies.
