Description
Prime Intellect is hiring an AI Infrastructure Engineer to build and operate its hosted training platform, which enables users to launch LoRA and full fine-tuning runs on managed GPU clusters. The role combines Kubernetes-based training and inference orchestration, Python control-plane agents, scheduling and autoscaling, GitOps, observability, FastAPI backend services, and Next.js/React/TypeScript developer-facing surfaces. It also involves interfacing with RL trainers, inference servers, and environment servers, and productizing new training capabilities. The position offers $150,000–$300,000 in cash compensation plus significant equity, remote or San Francisco work, visa sponsorship, and relocation support.

