Description
Joblogic is hiring an AI Evaluation Engineer to own quality for model-based AI agents across email, voice, SMS, WhatsApp, and CRM channels. The role defines rubrics and datasets, builds and calibrates LLM judges, runs offline and online evaluations, gates releases, analyzes production failures, evaluates classical ML models, and conducts adversarial and human-in-the-loop testing. It requires at least three years of model-evaluation experience, strong Python engineering, LLM and agent-evaluation expertise, and knowledge of RAG, observability platforms, statistics, and AI safety. The role is hybrid, with the possibility of working onsite with the UK team.
