Description
The Inference Stack Engineer will design and build components of an AI inference stack, including a Python-based DSL, compiler infrastructure, IR transformations, graph lowering, backend code generation, runtime systems, and workload profiling. The role focuses on optimizing model execution for latency, throughput, memory efficiency, and numerical stability, and requires strong C++ and Python software engineering, compiler or performance-critical systems experience, and knowledge of AI model execution and compute graphs.
