Description
The client’s Cloud Operations team is hiring a Site Reliability Engineer to keep user-facing services and production systems reliable, scalable, and secure. The role combines operational response, infrastructure automation, monitoring, deployment improvement, production debugging, capacity planning, and software development. The engineer will work with Linux and Windows systems, Ansible, Puppet, Terraform, Kubernetes, Docker, Nginx, HAProxy, Prometheus, public cloud platforms, and programming languages such as Python, Java, Golang, or Node.js, including AWS-to-Kubernetes migration and defining SRE KPIs.

