Description
ToolsGroup is hiring an experienced Site Reliability Engineer to lead IT service operations and a small IT/Ops team while serving as the senior technical escalation point for production services. The role focuses on diagnosing distributed-system failures, leading major incidents, troubleshooting Windows, Linux, Kubernetes, cloud, networking, identity, database, and API systems, automating operational work, improving observability and reliability, and translating technical decisions into business priorities. Candidates should have at least five years of production engineering experience, direct technical ownership of high-severity incidents, and hands-on experience with cloud, containers, observability, automation, and infrastructure as code. The position offers a salary of 55,000–68,000 per year plus a 10% bonus.
