Description
NVIDIA is hiring a Senior Inference Engineer focused on GPU kernel optimization for LLM inference. The role develops silicon-measured kernel benchmarking infrastructure, model-level performance analysis, and agentic kernel-optimization systems, then collaborates with compiler, hardware, kernel, and framework teams to improve production inference performance. Candidates need a master's or PhD in a relevant field or equivalent experience, at least six years of industry experience, expertise in Python and C++, GPU profiling tools, LLM inference frameworks, and GPU kernel optimization techniques.
