Description
The GPU Systems Engineer will lead the design and optimization of GPU inference runtimes, kernel dispatchers, memory planners, and distributed inference systems for Stable Diffusion, multimodal transformers, and video generation models. The role involves investigating cross-GPU bottlenecks, establishing GPU optimization standards and tooling, collaborating with research and hardware vendors, and mentoring engineers. Candidates need at least five years of experience in high-performance computing, GPU runtime systems, or ML infrastructure, along with expertise in CUDA, Triton, C++, distributed systems, and GPU interconnects.

