Description
fal is hiring a Senior Site Reliability Engineer to own the reliability and availability of customer-facing systems at scale. The role manages Kubernetes infrastructure, CI/CD and deployment pipelines, networking, service mesh, SLOs, incident response, observability, automation, and reliability improvements. It requires 5+ years of experience with production systems, Kubernetes, infrastructure-as-code, Linux networking, GitOps, Python or Go, monitoring, and AI/ML workloads, and is based in downtown San Francisco with health, dental, and vision insurance.
