Description
The employer is hiring an experienced HPC infrastructure engineer to lead the bringup, administration, and operations of a large-scale GPU training cluster. The role bridges researchers and bare-metal GPU hardware, ensuring SLURM jobs, parallel filesystems, networking, and anime-model training run reliably. Responsibilities include managing modern HPC software stacks such as SLURM, Kubernetes, Warewulf, MAAS, Ansible, Weka, VAST, Ceph, Tailscale, Grafana, and Prometheus, along with traditional Linux system administration. The position is based in Tokyo or San Francisco, with Bay Area preference and available visa sponsorship.
