Description
Fluidstack is hiring a Production Engineer to own reliability for named customer workloads, debug distributed systems across hardware, fabric, and scheduler layers, communicate incident updates to customers, and drive engineering fixes for recurring issues. The role requires at least five years of applied industry experience supporting large-scale compute customers, methodical distributed-systems debugging, and customer-facing incident communication; GPU training workloads, InfiniBand or RoCE, Slurm or Kubernetes, and NCCL debugging are bonuses.
