Skip to main content

HPC Infrastructure Engineer at AshbyHQ

Department: Research Team

Language
Setup
On-site
Location
Antioquia
Type
Full-time
Level
senior
Posted

Description

The employer is hiring an experienced HPC infrastructure engineer to lead the bringup, administration, and operations of a large-scale GPU training cluster. The role bridges researchers and bare-metal GPU hardware, ensuring SLURM jobs, parallel filesystems, networking, and anime-model training run reliably. Responsibilities include managing modern HPC software stacks such as SLURM, Kubernetes, Warewulf, MAAS, Ansible, Weka, VAST, Ceph, Tailscale, Grafana, and Prometheus, along with traditional Linux system administration. The position is based in Tokyo or San Francisco, with Bay Area preference and available visa sponsorship.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation