Description
Pathway is hiring a full-time ML Evaluation and Benchmarking professional to design rigorous benchmarks, define dataset standards, and build evaluation infrastructure for its post-transformer AI models. The role involves curating and assessing public and client-driven benchmarks, validating benchmark setups with baseline models, preparing benchmark-ready packages for R&D, maintaining documentation and shared evaluation terminology, tracking results and leaderboards, and contributing to customer-facing demos and proof points. Candidates should have ML/LLM evaluation, data science, or technical product experience and meet at least one distinction such as a top-tier ML publication, major LLM training contribution, experience at a leading ML research center, or elite programming/academic competition achievement.
