Description
SpaceX is hiring a Senior Site Reliability Engineer focused on Starshield’s software and GPU infrastructure. The role manages GPU and CPU deployments in Top Secret data centers, supports GPU-as-a-service, designs and productizes AI clusters at 100k+ GPU scale, automates on-premise Kubernetes and AI clusters, manages databases and distributed storage, and collaborates with AI engineers. The position requires a relevant bachelor’s degree and 5+ years of Linux experience, or 7+ years of software, DevOps, or site reliability experience; 5+ years of Kubernetes experience; and experience with Terraform, Ansible, containerization, scripting, and Python, C++, or Go. The role requires a Top Secret security clearance, extended hours, domestic and global travel, and ITAR eligibility. Compensation is $165,000–$265,000 for Level 3, with medical, vision, dental, retirement, and other benefits.
