Description
Fluidstack is hiring a Production Engineer to own the health, reliability, automation, and operations of large-scale GPU compute fleets. The role covers GPU fleet monitoring, repair and RMA automation, GPU qualification and burn-in, Redfish and BMC telemetry, Kubernetes-orchestrated bare-metal infrastructure, and incident response. Candidates should have hardware and firmware reasoning skills, experience shipping production automation, and familiarity with AI tooling; hardware lifecycle, RMA, BMC/Redfish, GPU qualification, workflow orchestration, metrics, alerting, and Go or Python are listed as bonuses. The position offers competitive salary and equity, retirement or pension benefits, health, dental, and vision insurance, and generous paid time off.
