Skip to main content

AI Evaluation & Benchmarking Engineer at Intergral

Department: EngineeringSeniority in posting: Mid-level

Compensation · Listed in posting

£60,000 – £75,000/yr

Setup
Remote - UK
Type
Full-time
Level
mid
Posted

Description

Intergral UK is hiring a full-time, remote AI Evaluation & Benchmarking Engineer in the United Kingdom to build automated evaluations and benchmarking systems for its AI-led observability platform, OpsPilot. The role covers agentic AI evaluation, synthetic workloads, ground truth, model and workflow benchmarking, OpenTelemetry-based environments, product telemetry, and continuous measurement of quality, reliability, latency, cost, and customer outcomes. The engineer will work independently with AI, software, platform, and SRE teams, build evaluation tooling, and help identify regressions and improvement opportunities. The position requires around three or more years of relevant technical experience, with no visa sponsorship.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation