Description
NVIDIA is hiring an RL Infrastructure Engineer to architect and build reinforcement-learning post-training infrastructure that scales from single-GPU experimentation to production across thousands of nodes. The role covers GPU, CPU, and LPU training-inference-rollout loops; fault tolerance, elastic scaling, and fast restarts; open-source RL framework contributions; and collaboration with hardware, networking, compiler, and researcher teams. Candidates need a master's or PhD in a relevant field, at least five years of experience in distributed systems, high-performance computing, deep-learning infrastructure, or ML systems engineering, and strong Python and C/C++ skills. The posting is marked hybrid and lists base salaries of 184,000–287,500 USD for Level 4 and 224,000–356,500 USD for Level 5.
