Description
The SRE – NOC role combines traditional Network Operations Center responsibilities with engineering-driven reliability practices. It operates in a 24x7 on-call rotation, leading incident response, monitoring service health, designing alerting strategies, reducing alert fatigue, automating operational tasks, and supporting Linux, cloud, Kubernetes, and production environments. The role requires Linux systems administration, incident management, production support, cloud and container experience, scripting, networking fundamentals, and familiarity with SLOs/SLIs; preferred qualifications include Terraform, Ansible, and security or compliance exposure. The position is an individual contributor reporting to the Manager, Network Operations and is marked #LI-Remote.
