Description
CloudLinux is seeking a highly experienced engineer to define Service Level Indicators (SLIs) and build a comprehensive monitoring system for Imunify360, a multi-layer Linux server security suite. The goal is to prevent silent failures like those that led to a 61-day security control disablement. This role involves designing and implementing a scalable telemetry pipeline, building alerting and escalation systems, and ensuring robustness across distributed systems. The ideal candidate will have substantial production-engineering or SRE experience, strong proficiency in Python, Go, or Rust, and deep understanding of time-series and event telemetry.
