Description
The Site Reliability Engineer will manage the full lifecycle of large-scale distributed services, from design and launch through monitoring, scaling, automation, and incident response. The role focuses on improving reliability, uptime, capacity, performance, and system health while using software development, algorithms, complexity analysis, and large-scale system design. It requires a bachelor’s degree or equivalent practical experience, five years of software development experience, three years of distributed-systems experience, and two years of project leadership and technical leadership experience.
