Description
The role leads the development and architecture of scalable, elastic, fault-tolerant distributed systems for large-scale data retrieval, storage, and processing. Responsibilities include defining scalability and reliability requirements, designing redundancy, replication, failover, load-shedding, throttling, rate-limiting, SLOs, telemetry, dashboards, alerts, fault-injection tests, data synchronization, security controls, compliance, infrastructure-as-code, and change-management automation. The role also handles production troubleshooting, incident response, operational readiness, mentoring, and technical project oversight. It is an IC4 career-level position with a U.S. hiring range of $114,600 to $234,600 per year and medical, dental, and vision insurance.
