Description
TensorWave is hiring a Hardware Diagnostics Engineer to perform GPU server burn-in and stress testing, triage hardware failures across GPUs, memory, drives, NICs, PSUs, and cabling, manage out-of-band server operations through IPMI and Redfish, apply firmware updates, drive vendor RMAs, maintain NetBox asset records, track failure patterns, improve runbooks, and automate repetitive work. The role requires 3–6 years of datacenter or infrastructure experience, enterprise server and out-of-band management expertise, strong Linux troubleshooting, scripting ability, and a methodical troubleshooting approach. Benefits include stock options, fully paid medical, dental, and vision insurance, health savings account contributions, disability insurance, retirement benefits, paid time off, holidays, parental leave, and other employee benefits.
