Description
NVIDIA is hiring an experienced Software Engineer for its DGX Cloud team to build and operate large-scale GPU clusters for AI workloads. The role focuses on Kubernetes-based GPU resource scheduling, monitoring, health management, incident response, and production reliability. Candidates need significant software engineering and Kubernetes experience, including cluster operations, operator development, node health monitoring, and GPU scheduling, along with systems programming in Go or Python and knowledge of distributed systems. The posting states that applications are accepted until October 3, 2026, and offers a base salary of 272,000 USD to 431,250 USD plus equity and benefits.
