Description
Thinking Machines is hiring an evergreen GPU Reliability Engineer to maintain the reliability of its GPU supercomputing fleet. The role owns the interface between hardware, firmware, and operating systems; investigates and resolves hardware, NIC, HBM, and kernel-driver issues; automates fleet monitoring and reliability analysis; manages firmware lifecycles; and engages GPU, server OEM, NIC, and storage vendors. The position is based in San Francisco, California, with an expected annual salary of $350,000–$475,000 USD and visa sponsorship.
