Description
Inferact is hiring a Head of Engineering to build and lead the organization developing systems that power vLLM and Inferact. The role requires an engineering leader with deep technical credibility at the inference layer, including GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design. The leader will partner with founders, scale a senior-heavy engineering team, recruit and develop ML systems talent, translate research and infrastructure work into execution plans, and improve engineering operations. The position is based in San Francisco, California, with relocation considered for exceptional candidates, and includes base compensation, equity, health, dental, vision, and a 401(k) company match.
