Skip to main content

HPC Site Reliability Engineer at Lambda

Department: Data Center Business

Compensation

$227,000 – $356,000/yr

Setup
Hybrid
Location
San Francisco, California
Level
senior
Posted

Description

Lambda is hiring an HPC Site Reliability Engineer to build and operate monitoring, alerting, and automation for large-scale AI HPC clusters. The role covers GPU, fabric, power, thermal, networking, and job-level observability; remote cluster deployment; infrastructure-as-code operations; runbooks and automated remediations; incident response; and troubleshooting of InfiniBand, RoCE, NCCL, GPU-direct, and related environments. Candidates need 7+ years of relevant experience, strong Linux and distributed-systems knowledge, Python and Go skills, and expertise with monitoring, automation, and configuration-management tools.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation