Description
The Site Reliability Engineer will combine software and systems engineering to build and operate large-scale, distributed, fault-tolerant systems for Google Cloud. Responsibilities include the full service lifecycle from design through deployment and refinement, pre-launch support, monitoring availability and latency, scaling systems through automation, and conducting incident response and blameless postmortems. The role requires a bachelor’s degree or equivalent practical experience, extensive software development, networking, distributed systems, and cross-team project leadership experience, along with strong problem-solving and communication skills.
