Description
Thinking Machines is hiring a Site Reliability Engineer to own the reliability, performance, and uptime of large-scale post-training and reinforcement learning training jobs. The role partners with research teams during active model runs, debugs failures across accelerators, networking, storage, schedulers, and training frameworks, builds monitoring and automated recovery, improves checkpointing and fault tolerance, develops internal tools, and participates in an on-call rotation. The position is based in San Francisco, California, with an expected annual salary of $300,000–$350,000 USD and visa sponsorship.
