Skip to main content

AI Evaluation Engineer at joblogic

Department: AI - Development - Joblogic

Setup
Hybrid
Location
Pakistan
Type
Full-time
Level
mid
Posted

Description

Joblogic is hiring an AI Evaluation Engineer to own quality for model-based AI agents across email, voice, SMS, WhatsApp, and CRM channels. The role defines rubrics and datasets, builds and calibrates LLM judges, runs offline and online evaluations, gates releases, analyzes production failures, evaluates classical ML models, and conducts adversarial and human-in-the-loop testing. It requires at least three years of model-evaluation experience, strong Python engineering, LLM and agent-evaluation expertise, and knowledge of RAG, observability platforms, statistics, and AI safety. The role is hybrid, with the possibility of working onsite with the UK team.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation