Summary from listing
Lam is hiring a Senior AIOps Reliability Engineer to build an AI-native operations layer for a hybrid enterprise estate spanning Azure, AWS, Google Cloud, on-premises data centers, colocation, manufacturing, and HPC facilities. The role designs and ships LLM-based agents, retrieval-augmented knowledge pipelines, machine-learning anomaly detection, automated incident response, safe remediation, AI safety and governance, and LLMOps/MLOps systems. It requires a computer science or engineering degree or equivalent experience, at least eight years of SRE, DevOps, infrastructure, observability, network engineering, or platform engineering experience, and production experience with LLM applications, cloud platforms, on-premises infrastructure, and networking. The role offers On-site Flex or Virtual Flex hybrid work models.
