Description
CoreWeave is seeking a curious, creative and persistent problem solver to join their Fleet Reliability Operations team. This team is responsible for the day-to-day provisioning, management and uptime of CoreWeave’s expanding fleet of server nodes. The individual will drive batches of server nodes through provisioning and validation processes while troubleshooting node or cluster problems, configure and maintain high-performance supercomputing clusters, troubleshoot hardware and software issues, monitor and analyze system performance, create and maintain documentation, and participate in on-call rotations.

