Description
Lambda is hiring a Senior Incident Manager to lead critical incident response across AI data center infrastructure, including GPU clusters, networking, storage, and cloud platforms. The role serves as Incident Commander during major outages, coordinates engineering, networking, facilities, and vendor teams, manages incident lifecycles and post-incident analysis, and improves operational resilience through runbooks, automation, observability, and reliability frameworks. Candidates need at least 8 years of experience in incident management, site reliability engineering, or infrastructure operations, along with expertise in large-scale distributed infrastructure and incident management frameworks. The position includes on-call rotation and offers health, dental, and vision coverage, wellness and commuter stipends for select roles, and a 401(k) plan with a 2% company match for USA employees.
