Description
FriendliAI is hiring a GPU Kernel Engineer to design, implement, and optimize low-level compute kernels for its large-scale GPU-accelerated AI inference platform. The role focuses on high-performance CUDA and C++ kernel development, reduced-precision and quantized inference, cross-vendor NVIDIA and AMD performance tuning, GPU libraries, and multi-modal model pipelines. Candidates need at least three years of GPU programming or performance-critical systems experience, a relevant bachelor’s or master’s degree, strong CUDA or ROCm/HIP expertise, and deep knowledge of GPU architecture.
