Description
fal is hiring a hands-on GPU Fleet Engineer to build and maintain software and tooling that manages thousands of bare-metal and cloud GPU servers. The role covers server provisioning, health monitoring, GPU diagnostics, automated recovery, metrics and alerting, OS and storage optimization, security hardening, and partner-driven incident resolution. Candidates should have at least three years of experience managing large server fleets, strong Python and Linux skills, infrastructure-as-code expertise, and familiarity with GPU and storage technologies. The position offers visa sponsorship, relocation support to San Francisco, and health, dental, and vision insurance.
