Description
Sarvam is hiring a Senior Performance Engineer, Kernels to own the kernel layer for its multi-node, multi-tenant GPU serving fleet. The role involves authoring and optimizing custom CUDA, DSL-based, and PTX kernels, including attention kernels, to improve production p99 performance. Candidates need at least five years of ML systems experience, including two years authoring production CUDA kernels, along with expertise in CUDA, CUTLASS/CuTe, PTX, Nsight Compute, and multi-architecture GPU systems.
