Description
Thinking Machines is hiring an Infrastructure Research Engineer to design, implement, and optimize distributed training systems for large-scale model training across thousands of GPUs and nodes. The role focuses on high-performance training, reusable frameworks, reliability, maintainability, security, and collaboration with researchers and engineers, with an emphasis on making experimentation and training fast and reliable. The position is based in San Francisco, California, offers visa sponsorship, and provides health, dental, vision, PTO, parental leave, and relocation benefits.
