Description
NVIDIA is hiring a Software Engineer to design, develop, and maintain engineering solutions for validating, monitoring, operating, and improving GPU clusters at scale for internal machine learning researchers. The role involves AIOps and Agentic AI research, on-call support, and self-service reliability and performance improvements, with requirements including a BS/MS in Computer Science, Engineering, or equivalent experience, at least two years of software or platform engineering experience, and expertise in Linux, Python, C++, Rust, Docker, Kubernetes, GitLab CI, and related ML infrastructure technologies.
