Summary from listing
The Site Reliability Engineer will ensure the availability, performance, reliability, and resilience of cloud-native platforms and customer-facing services. The role owns observability, SLIs/SLOs/SLAs, incident response, root cause analysis, performance optimization, reliability engineering practices, automation, CI/CD, infrastructure automation, and production on-call support. It requires experience with high-availability environments, Linux, Docker and Kubernetes, observability tools, CI/CD pipelines, Infrastructure as Code, networking, cloud-native architectures, and distributed systems.
