Description
fal is hiring a seasoned Site Reliability Engineer to own and operate customer-facing production infrastructure at scale, with a focus on Kubernetes, CI/CD, networking, observability, incident response, and automation. The role emphasizes improving reliability and availability across the stack, building resilient deployment and monitoring systems, and using automation and AI to accelerate issue resolution and software delivery. The position is based in San Francisco with remote flexibility considered for senior and staff levels, and includes relocation assistance plus medical, dental, and vision benefits.
