Description
Kraken is hiring an AI Compute and Infrastructure Engineer to own and operate GPU and accelerator clusters for model training, inference, evaluation, and experimentation. The role covers cluster scheduling, orchestration, placement, quota management, observability, reliability, incident response, cost optimization, and integration of serving frameworks and hardware. Candidates need at least five years of infrastructure engineering experience, hands-on GPU cluster operations, strong systems and Python skills, and familiarity with ML serving and distributed systems.
