Skip to main content

Reliability Engineer, Supercomputing at ThinkingMachines

Department: Core Engineering

Compensation

$350,000 – $475,000/yr

Setup
On-site
Location
San Francisco, California
Type
Full-time
Level
senior
Posted

Description

Thinking Machines is hiring an evergreen GPU Reliability Engineer to maintain the reliability of its GPU supercomputing fleet. The role owns the interface between hardware, firmware, and operating systems; investigates and resolves hardware, NIC, HBM, and kernel-driver issues; automates fleet monitoring and reliability analysis; manages firmware lifecycles; and engages GPU, server OEM, NIC, and storage vendors. The position is based in San Francisco, California, with an expected annual salary of $350,000–$475,000 USD and visa sponsorship.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation