Description
Microsoft is hiring a Senior Supercomputing Operations Engineer to own the reliability and operations of GPU interconnect fabrics and large-scale supercomputing clusters supporting AI training and other compute-intensive workloads. The role leads incident triage, root-cause analysis, cross-stack debugging, operational playbooks, and automation across hardware, firmware, drivers, and software. Candidates need a bachelor's degree in computer science or a related technical field plus at least four years of technical engineering experience with coding; equivalent experience is accepted, while the preferred qualifications include a master's degree, additional experience, and production HPC or interconnect-fabric expertise. The typical U.S. base pay range is USD 119,800–234,700 per year, with higher ranges for the San Francisco Bay area and New York City metropolitan area.
