Description
Recorded Future is hiring an experienced Site Reliability Engineer to improve reliability, scalability, performance, security, and operational excellence for critical systems. The role focuses on AWS infrastructure, observability, automation, incident response, and collaboration with engineering teams to support high availability and resilient applications. The position requires at least 3 years of relevant experience, strong Linux and troubleshooting skills, and hands-on work with Terraform, Chef, Grafana, ELK, and Prometheus, with preferred experience in Kubernetes, Kafka, RabbitMQ, MongoDB, OpenTelemetry, CI/CD, and distributed systems.
