Description
Hyperbolic Labs is hiring a Site Reliability Engineer to maintain the reliability, performance, security, and economic efficiency of its GPU marketplace and AI infrastructure. The role covers SLO and SLA management, capacity planning, incident response, monitoring and alerting, deployment automation, tenant and workload isolation, key management, compliance, and infrastructure hardening. Preferred qualifications include GPU infrastructure experience, distributed systems knowledge, multi-tenancy and container security, chaos engineering, cloud cost optimization, and experience with high-uptime systems.
