Description
Sciforium is hiring a GPU Kernel Engineer to design, implement, and optimize custom GPU kernels for large-scale AI systems, including LLM training and inference. The role spans low-level GPU programming with C++, PTX, CUDA, ROCm, Triton, and JAX Pallas; performance profiling and optimization; integration of kernels into PyTorch, JAX, and custom runtimes; and collaboration with researchers, distributed systems, model-serving, and hardware-vendor teams. Candidates need at least five years of GPU kernel or high-performance computing experience, a relevant advanced degree, and strong expertise in GPU memory models and performance optimization.
