Description
CMT is hiring a Principal Site Reliability Engineer I, Machine Learning to own the operational health, observability, scalability, security, cost, and reliability of Ray clusters on AWS EKS and Databricks workloads on AWS EC2. The role includes infrastructure-as-code with Terraform and CI/CD, incident response, postmortems, on-call participation, and maintenance of EC2 and EKS infrastructure. It is a hybrid position based in Cambridge, Massachusetts, with two days per week onsite, and offers a base salary of $142,000 to $177,600 plus bonus, equity, and benefits.
