Description
Bugcrowd is hiring a Reinforcement Learning Engineer specializing in Reinforcement Learning from Verifiable Rewards (RLVR) to build and scale automated verification environments that convert software vulnerabilities into deterministic reward signals for LLM reasoning models. The role involves developing test harnesses, sandboxes, execution environments, telemetry, instrumentation, and scalable distributed training infrastructure, while collaborating with frontier AI researchers and benchmarking cybersecurity capabilities. The position is fully remote and offers a base salary of $176,400 to $242,550.
