Description
Prime Intellect is hiring a Distributed Training Systems Engineer to build and optimize the distributed training infrastructure for pre-training and large-scale reinforcement-learning workloads. The role focuses on improving training efficiency across compute, memory, networking, and scheduling layers; designing low-level kernel, communication, and runtime optimizations; supporting data, tensor, and pipeline parallel workloads; and contributing to RL training stacks and open-source infrastructure. Candidates should have strong AI/ML systems experience, familiarity with PyTorch and distributed-training frameworks, GPU profiling expertise, and experience with large-scale parallel training techniques. The position offers $150,000–$350,000 in cash compensation plus equity, flexible remote or in-person work in San Francisco, and visa sponsorship and relocation assistance.

