Description
Sarvam is hiring a GPU Fleet Reliability Engineer to operate and maintain a large multi-vendor GPU fleet supporting long-running distributed training and production inference. The role covers fleet provisioning, observability, capacity, health, on-call ownership, runbooks, postmortems, and internal tooling, with candidates expected to bring depth in one of five areas: distributed high-performance storage, fabric and RDMA networking, GPU systems reliability, Kubernetes platform reliability, or training and inference workload reliability. The posting requires at least five years of infrastructure or site reliability engineering experience, including two years operating GPU clusters at scale, along with Python or Go proficiency and cross-area fluency.
