Description
Pathway is hiring a full-time, permanent machine learning evaluation and benchmarking professional to design rigorous benchmarks, define dataset standards, and build evaluation infrastructure for its post-transformer AI models. The role involves curating and validating public and client-driven benchmarks, running baseline models, preparing benchmark-ready packages for R&D, maintaining documentation and shared evaluation terminology, tracking results and leaderboards, and contributing to customer demos and public proof points. Candidates should have ML/LLM evaluation, data science, or technical product experience and meet at least one advanced achievement criterion, such as a leading ML publication, major LLM training contribution, experience at a leading ML research center, or elite programming/academic competition distinction. The position is remote, with optional access to offices in Palo Alto, Paris, or Wroclaw, and candidates in the EU, United Kingdom, United States, and Canada are considered.
