Description
Kog is hiring a GPU Engineer to work on low-level execution and GPU performance for its inference engine. The role focuses on a monokernel decode pipeline, AMD and NVIDIA kernel optimization, memory-bound execution, profiling, inter-GPU communication, and scaling third-party MoE models. Candidates should have experience below the framework layer with hardware behavior or performance, including CUDA, HIP, PTX, CDNA ISA, kernels, profiling, or related systems work. The role is based in Paris with at least one week per month of travel and accommodation support for candidates outside the region.
