Description
The role is responsible for designing, operating, and optimizing NEURA’s large-scale AWS HyperPod GPU cluster infrastructure supporting foundation model training and customer fine-tuning workloads. Responsibilities include configuring HyperPod/Slurm and HyperPod/EKS orchestration, improving cluster stability and fault tolerance, managing workload priorities and GPU utilization, building self-service tooling, developing user documentation, and negotiating AWS capacity and cost strategies. The position requires at least five years of infrastructure or systems engineering experience focused on GPU clusters or HPC operations, along with hands-on AWS HyperPod experience, knowledge of Slurm and Kubernetes, distributed-training expertise, and cloud cost-management skills.

