Description
Replay is hiring its first Data Scientist focused on evaluation to establish decision-grade measures for the quality and safety of de-identified enterprise data. The role builds evaluation corpora, experiments, quality metrics, pipelines, and feedback loops; analyzes failure patterns across models, prompts, judges, thresholds, and workflows; and connects offline measures to escaped sensitive information, utility loss, review burden, and delivery decisions. It requires hands-on Python and SQL, statistical evaluation knowledge, startup ownership, and the ability to independently verify AI-generated evidence.
