Description
Lightly AG is hiring a remote, part-time AI Evaluation Researcher to evaluate agentic AI-generated peer reviews of machine learning and artificial intelligence research papers against expert human reviews. The role involves reading and understanding papers, establishing human-review baselines, scoring AI reviews against a structured rubric, identifying hallucinations and missed technical issues, comparing multiple AI reviews, verifying academic literature, and providing evidence-based rationales. Candidates should have a master’s, PhD, or graduate-level background in a relevant technical field, research-paper experience, familiarity with major ML/AI venues, prior academic peer-review experience, strong analytical and written communication skills, and attention to detail.

