Skip to main content

Senior Cluster Site Reliability Engineer at The Voleon Group

Department: Software

Compensation

$205,000 – $235,000/yr

Setup
Remote
Location
Berkeley, California
Type
Full-time
Level
senior
Posted

Description

Voleon is hiring a Senior Cluster Site Reliability Engineer to scale and operate research compute clusters supporting machine learning research. The role focuses on maintaining high uptime and reliability, responding to outages, diagnosing recurring issues, improving observability, enforcing fair cluster usage, forecasting growth, and optimizing cost and usability across on-premises and cloud infrastructure. Candidates need at least five years of SRE or DevOps experience, knowledge of HPC and batch compute frameworks, scripting and infrastructure-as-code skills, cloud experience, distributed storage knowledge, and a bachelor's degree in computer science.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation