Description
Scale is hiring a Research Scientist to join its GenAI Research Organization's evaluation pod. The role focuses on building benchmarks and diagnostic methods for frontier LLMs and agents, analyzing text and multimodal model failure modes, connecting failures to post-training techniques such as SFT, RLHF, and reward modeling, and publishing research findings. The position requires expertise in deep learning, reinforcement learning, large-scale model fine-tuning, LLM evaluation, and benchmark development, with a Ph.D. or Master's degree preferred. Compensation for eligible full-time roles in San Francisco, New York, and Seattle is $165,600–$207,000 USD, with equity and benefits.
