Description
The Lead Site Reliability Engineer will be responsible for implementing and managing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts. They will develop resilient systems, lead incident response and post-mortem analyses, automate incident detection and response, and design full observability using modern tools. The role involves collaborating with development and operations teams, championing Infrastructure as Code, participating in chaos engineering initiatives, and driving advanced alerting.
