Description
Microsoft is hiring a Site Reliability Engineer for its Azure Specialized AI Infrastructure team in India. The role focuses on the reliability, scalability, security, performance, and automation of large-scale distributed systems supporting HPC and generative AI workloads, including incident response, root cause analysis, cloud and AI infrastructure technologies, and specialized hardware such as GPUs and InfiniBand. The position requires substantial professional software engineering experience, service operations and reliability experience, and experience with cloud or distributed systems; applications are accepted until the position is filled.
