Description
The Tools & Platforms Site Reliability Engineer owns the reliability, observability, automation, and continuous improvement of infrastructure tooling platforms across Splunk, Terraform, CI/CD, ITSM, and AI-augmented operations. The role defines SLIs, SLOs, and error budgets; engineers auto-remediation and observability-as-code; monitors model APIs, RAG platforms, and agent workflows; and leads post-mortems in a regulated financial services environment. It requires at least five years of Splunk Enterprise Architecture and Design experience, 15 years of full-time education, and specified certifications and technical skills.
