Description
Cosm is hiring a Site Reliability Engineer to design, implement, automate, and maintain the technology infrastructure supporting its operations center. The role focuses on monitoring and alerting, application and infrastructure deployment, incident management, documentation, training, and collaboration with product, engineering, operations, security, and business teams. The engineer will work with hybrid systems, AWS and Azure, Kubernetes, Terraform, Linux and Windows Server, and observability technologies, while providing technical guidance and mentoring junior team members.
