Description
Prime Intellect is hiring an AI Infrastructure Engineer to build and optimize systems for large-scale reinforcement learning and distributed model training. The role focuses on improving training efficiency across compute, memory, networking, and scheduling layers; designing low-level kernel, communication, and runtime optimizations; supporting data, tensor, and pipeline parallel workloads; and contributing to RL training stacks and open-source infrastructure. Candidates should have strong AI/ML systems experience, familiarity with PyTorch and distributed training frameworks, GPU profiling expertise, and experience optimizing training performance. The position offers $150,000–$350,000 in cash compensation plus equity, flexible remote or in-person work from San Francisco, and visa sponsorship and relocation support.

