Description
Lambda is hiring an HPC Site Reliability Engineer to build and operate monitoring, alerting, and automation for large-scale AI HPC clusters. The role covers GPU, fabric, power, thermal, networking, and job-level observability; remote cluster deployment; infrastructure-as-code operations; runbooks and automated remediations; incident response; and troubleshooting of InfiniBand, RoCE, NCCL, GPU-direct, and related environments. Candidates need 7+ years of relevant experience, strong Linux and distributed-systems knowledge, Python and Go skills, and expertise with monitoring, automation, and configuration-management tools.
