Description
ai& is hiring an Inference & Serving Engineer to build and optimize a high-performance, multi-tenant inference serving stack across heterogeneous hardware. The role focuses on selecting and tuning inference frameworks, improving latency and throughput through techniques such as disaggregated prefill/decode, speculative decoding, and continuous batching, and scaling architectures from small clusters to multi-node deployments. Responsibilities also include memory and KV-cache optimization, Day 0 model support, cross-stack integration, production debugging, and hands-on technical leadership. The position requires deep inference-engine experience, knowledge of multimodal generative AI, and a track record of distributed-systems engineering.
