Skip to main content

GPU Cluster Engineer (human) at neura-robotics-gmbh

Department: Software Engineering

Language
Setup
On-site
Location
Riederich, Baden-Württemberg
Level
senior
Posted

Description

The role is responsible for designing, operating, and optimizing NEURA’s large-scale AWS HyperPod GPU cluster infrastructure supporting foundation model training and customer fine-tuning workloads. Responsibilities include configuring HyperPod/Slurm and HyperPod/EKS orchestration, improving cluster stability and fault tolerance, managing workload priorities and GPU utilization, building self-service tooling, developing user documentation, and negotiating AWS capacity and cost strategies. The position requires at least five years of infrastructure or systems engineering experience focused on GPU clusters or HPC operations, along with hands-on AWS HyperPod experience, knowledge of Slurm and Kubernetes, distributed-training expertise, and cloud cost-management skills.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation