Description
NVIDIA is hiring an experienced Kubernetes Software Engineer to help scale its AI infrastructure by building and operating large GPU clusters. The role focuses on custom software for GPU resource scheduling, monitoring and health management, reliability and availability, incident management, and production AI cluster performance. Candidates need significant software engineering and Kubernetes experience, including cluster operations, operator development, node health monitoring, and GPU scheduling, along with systems programming in Go or Python and knowledge of distributed systems. The posting accepts applications through October 2, 2026, and lists base salary ranges for Level 3 and Level 4 employees.
