Description
Thinking Machines is hiring an evergreen GPU Supercomputing Engineer to design, build, and operate GPU clusters supporting large-scale model training and inference. The role covers cluster provisioning, imaging, capacity planning, software abstraction, Kubernetes or Slurm scheduling, performance monitoring, storage and artifact management, and collaboration with researchers. Candidates need a bachelor’s degree or equivalent experience, backend programming ability, and experience with large-scale clusters and container orchestration. The position is based in San Francisco, California, with an expected annual salary of $350,000–$475,000 USD, visa sponsorship, and health, dental, vision, PTO, parental leave, and relocation benefits.
