Description
Inferact is seeking an Infrastructure Engineer to develop scalable distributed systems for powering global-scale AI inference using vLLM. The role involves designing and implementing foundational layers to ensure low-latency and high-reliability service of models across numerous accelerators. The goal is to simplify deploying frontier models at scale, akin to serverless databases.
