Description
The Senior Site Reliability Engineer will lead operational reliability, observability, incident management, production stability, and cost optimization across a complex multi-cloud platform. Responsibilities include building a unified metrics, logs, and traces architecture; establishing SLI/SLOs and error budgets; optimizing AWS and GCP costs; improving Kubernetes clusters and recovery mechanisms; and developing incident and deployment frameworks. The role requires at least five years of DevOps, SRE, or cloud-platform experience, technical leadership experience, production distributed-systems experience, deep Kubernetes knowledge, and experience establishing observability solutions. The position is based in Gothenburg, with English required and Mandarin an advantage.
