Description
Anthropic is hiring a Research Engineer to build, own, and improve the core reinforcement learning training system used for production and research model training. The role spans orchestration, environments, training, inference, evaluation, experimentation, optimization, debugging, and API design, with a focus on scaling RL training and collaborating with research teams. Candidates need Python, large-scale machine learning training, modern ML framework, experimental design, and communication skills; experience with reinforcement learning for large language models, distributed training, LLM architectures, profiling, RL environments, Rust, or C++ is preferred. The role offers a $500,000–$850,000 USD annual salary, requires a bachelor’s degree or equivalent, follows a hybrid policy with at least 25% office time, and may offer visa sponsorship.

