Description
The role operates and maintains a Kubernetes platform, managing clusters and nodes, persistent storage, reliability and SLOs, observability, monitoring, incident management, on-call operations, automation, runbooks, and CI/CD workflows. It requires at least two to three years of Linux systems operations experience, practical Kubernetes and Docker experience, monitoring or logging-stack knowledge, and familiarity with CI/CD and Git-based processes. The position involves collaboration with development, operations, and business teams, participation in on-call rotations, and contribution to platform reliability improvements.
