Skip to main content

Senior Site Reliability Engineer, DGX Cloud at NVIDIA

Department: Site Reliability Engineering

Language
Setup
Remote - China
Location
Zurich, Zurich
Level
senior
Posted

Description

NVIDIA is hiring a Senior Site Reliability Engineer to maintain and improve large-scale Kubernetes clusters supporting DGX Cloud for AI researchers and enterprise clients. The role covers operational reliability, SLO/SLI and error-budget management, service launch and post-launch monitoring, GPU workload optimization across major cloud providers, automation, incident triage, root-cause analysis, and on-call support. Candidates need a BS in Computer Science or related technical field or equivalent experience, at least 10 years of production-service operations experience, and expertise in Kubernetes, containerization, microservices, infrastructure automation, programming, Linux, networking, cloud security, observability, and SRE practices.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation