Description
ClickHouse is hiring a Site Reliability Engineer to build and lead reliability processes for its cloud infrastructure supporting ClickHouse databases. The role focuses on scalable, secure, highly available, and fault-tolerant distributed systems; SLOs and SLAs; monitoring, alerting, incident response, post-mortem analysis, chaos initiatives, on-call processes, and software platforms that improve operational efficiency. Candidates need a bachelor’s or master’s degree in computer science or a related field, at least eight years of site reliability engineering or related experience, production ClickHouse experience, and hands-on experience with Go or Python, cloud platforms, distributed databases, container orchestration, and automation tools.
