Description
CoreWeave is hiring a Fleet Reliability Operations Engineer to provision, configure, maintain, monitor, and troubleshoot large-scale GPU supercomputing clusters and their networking and platform dependencies. The role involves Linux system administration, hardware and software troubleshooting, observability analysis, documentation, process improvement, and on-call rotations. Candidates should have strong Linux and scripting skills; data center, Kubernetes, HPC, and observability experience are preferred. The position includes medical, dental, pension, life assurance, critical illness, and other benefits.
