Description
Callosum is hiring an AI Infrastructure Performance Engineer to bridge internal engineering with real-world performance. The role runs heterogeneous model-inference experiments, develops reproducible deployment patterns, improves caching, scheduling, batching, and resource allocation, and provides performance feedback to accelerator, engine, and infrastructure teams. It also builds benchmarking harnesses, regression suites, and dashboards. The position requires production or production-equivalent inference experience, multi-node GPU familiarity, strong performance-characterisation skills, and knowledge of serving frameworks such as Dynamo or Triton. It is based in Callosum’s London office and offers visa sponsorship and relocation benefits.
