Description
Voleon is hiring a Senior Cluster Site Reliability Engineer to scale and operate research compute clusters supporting machine learning research. The role focuses on maintaining high uptime and reliability, responding to outages, diagnosing recurring issues, improving observability, enforcing fair cluster usage, forecasting growth, and optimizing cost and usability across on-premises and cloud infrastructure. Candidates need at least five years of SRE or DevOps experience, knowledge of HPC and batch compute frameworks, scripting and infrastructure-as-code skills, cloud experience, distributed storage knowledge, and a bachelor's degree in computer science.
