Description
IDeaS is seeking an experienced Site Reliability Engineer (SRE) with over five years of hands-on experience in AWS, Kubernetes, Terraform, SQL Server, Datadog, and Java-based service reliability. The role focuses on ensuring the stability, performance, and reliability of the Hospitality Revenue Management System (RMS), a multi-tenant SaaS platform for hotels worldwide. Key responsibilities include defining SLIs/SLOs, improving Java microservice reliability, optimizing AWS services like EC2, EKS, S3, RDS, Lambda, and API Gateway, implementing Infrastructure-as-Code with Terraform, and configuring Datadog for comprehensive observability across metrics, logs, traces, and JVM profiling. This role demands strong collaboration with engineering to resolve critical issues and enhance overall system resilience.

