Description
The role is an AI-safety research engineer focused on testing and improving reinforcement-learning agents in Frontier labs. Responsibilities include running experiments on reward hacking, cheating, unsafe tool use, deception, and long-horizon behavior; breaking and closing graders; building trajectory monitors, red-team harnesses, and safety benchmarks; securing environments against network escapes and credential misuse; and conducting research on oversight, control, or long-horizon autonomy. The position requires strong engineering, LLM-safety, red-teaming, monitoring, and production-evaluation skills, with a base salary of $250,000–$450,000, a performance bonus, equity, and health, dental, vision, and flexible PTO benefits.
