Description
Mirendil is hiring an engineer to build and improve the post-training infrastructure for frontier reasoning models. The role combines research and infrastructure work, focusing on large-scale reinforcement-learning training, performance optimization, evaluation and benchmarking, data collection and feedback pipelines, and collaboration with multiple teams to bring RL experiments into production training runs.
