Description
Fluidstack is hiring a Production Engineer to own the health, reliability, automation, and operation of large-scale GPU compute fleets. The role covers GPU fleet monitoring, repair and RMA automation, GPU qualification and burn-in, Redfish and BMC telemetry, Kubernetes-orchestrated bare-metal operations, and incident response. Candidates should have at least five years of applied industry experience, hardware and firmware reasoning skills, production automation experience, and familiarity with AI tooling; hardware lifecycle, RMA, BMC/Redfish, GPU qualification, workflow orchestration, metrics, alerting, and Go or Python are bonuses. The position offers competitive salary and equity, retirement or pension benefits, health, dental, and vision insurance, and generous paid time off.
