Description
The role builds and calibrates LLM-based judges for Moveworks’ agentic AI evaluation platform, including rubrics, confidence reporting, human-label calibration, small-model fine-tuning, and guardrails against correlated blind spots. It also develops a process-reward signal to optimize agent prompts, tool selection, planning, retrieval, and routing in simulated environments. The position requires 8+ years of applied ML, data science, or ML-adjacent engineering experience, strong Python, and expertise in LLM evaluation, human annotation, reward modeling, or agent trajectory analysis.
