Description
The role focuses on building and optimizing production LLM serving and inference systems, improving GPU and CPU performance, addressing KV cache, memory, storage, and throughput bottlenecks, and designing scalable systems for RAG and retrieval-heavy AI workloads. It seeks engineers with meaningful production AI-systems experience, deep systems-layer knowledge, and the ability to work across architecture and implementation; a PhD is preferred but secondary to real-world systems experience.
