Description
Fluidstack is hiring a Production Engineer to own the health, reliability, automation, and operations of large-scale GPU compute fleets. The role covers GPU fleet monitoring, repair and RMA automation, GPU qualification and burn-in, Redfish and BMC telemetry, Kubernetes-orchestrated bare-metal infrastructure, incident response, and scalable observability. Candidates should have hardware and firmware reasoning, production automation experience, AI tooling fluency, and strong incident discipline; hardware lifecycle, BMC/Redfish, GPU qualification, workflow orchestration, metrics, and Go or Python are bonuses. Compensation includes salary, equity, retirement or pension, health, dental, and vision insurance, and generous PTO.
