Description
The Infrastructure Engineer will design, deploy, and operate large distributed GPU clusters for training, evaluation, and serving AI workloads. The role extends Kubernetes and Slurm scheduling and orchestration, builds self-serve cluster-management software, manages storage and artifact paths, improves reliability and observability, and partners with researchers to optimize large-scale runs. Candidates should have experience with GPU clusters and container orchestration, strong systems knowledge, cloud platform familiarity, and knowledge of CUDA, NCCL, and distributed-workload profiling.
.png&sig=GmGZl0hIEC4KeSLZOLIDYB_qjM5Xym0OKo-pzrtoUqQ)
