Description
OpenAI is seeking a Process Management Engineer specializing in Training Runtime to design and develop the core distributed runtime for large-scale machine learning workloads. The role involves creating robust, scalable, and high-performance components using Rust, focusing on maximizing researcher productivity and hardware utilization. The engineer will work on orchestrating and monitoring machine learning workloads across massive supercomputers, improving reliability, observability, and fault tolerance, and debugging complex distributed systems. This hybrid role is based in London, UK, requiring proficiency in Python and/or another systems programming language like C++, along with strong Linux knowledge and experience with asynchronous/concurrent systems.
