Description
The Site Reliability Engineer will protect and improve the availability, performance, and resilience of systems supporting bet365’s global online gambling products. The role combines software engineering, automation, observability, incident response, and reliability engineering across development, IT Operations, and SRE teams. Responsibilities include building resilient tools and operational APIs, automating system management, developing telemetry and instrumentation, configuring Cloudflare edge services, diagnosing distributed-system incidents, participating in incident response and root-cause analysis, and improving reliability and observability across teams. The position requires Python, Golang, JavaScript, or similar software engineering experience, observability and infrastructure-as-code expertise, and is eligible for the company’s hybrid work-from-home policy.
