Description
The employer is hiring a Senior DevOps Engineer to own the reliability, observability, capacity, disaster recovery, incident response, and operational automation of a large distributed video-monitoring and AI-alerting platform. The role covers GPU-backed computer-vision inference, Python and legacy Java services, AWS infrastructure, databases, messaging, and a remote edge fleet, with responsibilities including SLI/SLO management, alert quality, incident command, Terraform, Linux administration, and Python/Golang/Bash automation. The position is permanent and hybrid, requires AWS certification and 10+ years of site reliability or production operations experience, and offers benefits.
