Skip to main content

Infrastructure SRE - HPC at Sarvam

Department: Infrastructure

Setup
On-site
Location
Bengaluru, Karnataka
Level
senior
Posted

Description

Sarvam is hiring a GPU Fleet Reliability Engineer to operate and maintain a large multi-vendor GPU fleet supporting long-running distributed training and production inference. The role covers fleet provisioning, observability, capacity, health, on-call ownership, runbooks, postmortems, and internal tooling, with candidates expected to bring depth in one of five areas: distributed high-performance storage, fabric and RDMA networking, GPU systems reliability, Kubernetes platform reliability, or training and inference workload reliability. The posting requires at least five years of infrastructure or site reliability engineering experience, including two years operating GPU clusters at scale, along with Python or Go proficiency and cross-area fluency.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation