Description
Lambda is hiring a Site Reliability Engineer to operate and scale its multi-tenant cloud networking platform and SDN infrastructure, manage Kubernetes-based control-plane and dataplane services on SmartNICs, and develop automation and reliability tooling. The role also covers observability, incident management, capacity planning, CI/CD and GitOps workflows, and on-call participation. Candidates need at least five years of experience in site reliability, production engineering, or a similar role, along with expertise in distributed systems, Kubernetes, Linux, networking, observability, and infrastructure automation. The position requires presence in the San Francisco, San Jose, or Bellevue office four days per week, with Tuesday as the designated work-from-home day.
