Description
Pluralis Research is hiring an RL Post-Training Engineer to build and ship decentralized post-training systems for large language models. The role owns the end-to-end RL training loop, including rollout ingestion, reward computation, policy updates, weight synchronization, asynchronous and high-latency algorithms, evaluation, and the first public release of a post-trained model. Candidates should have hands-on experience with RL post-training, production Python and PyTorch, and distributed or asynchronous systems; the position is remote-first, offers equity-heavy compensation, and provides optional visa sponsorship to Australia or the United States.
