Description
CloudLinux is hiring a remote-first production engineering/SRE specialist to establish an SLI/SLO, telemetry collection, alerting, escalation, and incident-management function for the Imunify360 Linux server security product. The role involves defining measurable health indicators for approximately 70 cloud-side and agent-side components, building privacy-conscious fleet telemetry in Python, Go, and Rust, implementing SLO-based burn-rate alerting and squad-owned escalation, and improving detection of silent security-control degradation from a 61-day baseline to under 24 hours.
