Description
The Site Reliability Engineer II will operate and improve a hybrid Kubernetes and on-premises infrastructure spanning RKE2/Rancher clusters, physical HPE data-center servers, and Azure Kubernetes Service. The role includes troubleshooting production systems, supporting disaster recovery, participating in on-call incident response, and reducing toil through automation, infrastructure as code, CI/CD, and observability improvements. Candidates need at least three years of SRE, platform, systems engineering, or DevOps experience, hands-on production Kubernetes and data-center infrastructure experience, strong Linux and networking fundamentals, and scripting or Git experience.
