Description
Seeking an experienced Site Reliability Engineer (SRE) with over five years of hands-on experience in AWS, Kubernetes, Terraform, SQL Server, Datadog, and Java-based service reliability. The role focuses on ensuring stability, performance, and reliability of a cloud-native Hospitality Revenue Management System (RMS), a multi-tenant SaaS platform for hotels worldwide. Key responsibilities include defining SLIs/SLOs, improving Java microservice reliability, optimizing AWS services like EC2, EKS, S3, RDS, Lambda, and implementing cost optimization strategies. Additionally, the engineer will utilize Infrastructure-as-Code (Terraform) for reusable module creation, enforce GitOps principles, and configure Datadog for comprehensive observability across metrics, logs, traces, JVM profiling, and API uptime monitoring. This role demands strong expertise in Java internals, SQL Server tuning, and proficiency in Python/Bash/Go scripting.

