Description
Sciforium is hiring a GPU Cluster Engineer to own the software stack of GPU clusters, from kernel and driver tuning through schedulers, containers, and ML frameworks. The role builds and validates production-ready node images, automates fleet upgrades and self-healing workflows, manages Kubernetes and Slurm/Run:AI workloads, maintains NVIDIA and AMD accelerator stacks, and troubleshoots distributed training and inference performance issues. Candidates need at least five years of systems or infrastructure engineering experience, a relevant bachelor's or master's degree, deep Linux and GPU-stack expertise, and experience with configuration management, provisioning, containers, and distributed filesystems.
