Description
SpaceX is hiring a Site Reliability Engineer focused on Starshield’s software and GPU infrastructure. The role manages GPU and CPU deployments in Top Secret datacenters, supports GPU-as-a-service, designs and productizes AI clusters at 100k+ GPU scale, automates on-premise Kubernetes and AI clusters, and maintains databases, monitoring, and distributed storage. The position requires a relevant bachelor’s degree or equivalent experience, Linux, infrastructure automation, containerization, scripting, and development experience, along with a Top Secret security clearance and willingness to work extended hours, weekends, and travel.
