Description
CloudFactory is hiring a Site Reliability Engineer to own the reliability, observability, developer tooling, automation, and secure-by-default infrastructure of its AI platform and core services. The role covers model serving and inference infrastructure, GPU-backed endpoints, autoscaling, SLOs, incident response, ML and LLM observability, GitOps, Terraform, and reusable platform components. Candidates need at least five years of infrastructure, DevOps, or SRE experience with Kubernetes, production operational experience, cloud and scripting skills, and strong cross-functional collaboration; AI platform experience is strongly preferred.
