Description
Prior Labs is hiring an Infrastructure Engineer to own and evolve multi-cluster GPU infrastructure for training tabular foundation models. The role covers Slurm and GCP cluster architecture, scheduling, reliability, cost optimization, GPU utilization, distributed-training performance, hardware selection, capacity planning, and developer-productivity tooling. Candidates should have at least three years of experience building and operating production GPU infrastructure or distributed training systems, strong Slurm and Python/PyTorch expertise, and a track record of improving training throughput or cost efficiency. The team is based in Berlin, Freiburg, and New York, with remote work possible in exceptional cases.

