Description
Intel is hiring an AI Infrastructure Engineer to optimize large language model inference on next-generation Intel GPUs. The role spans profiling and resolving cross-stack bottlenecks, developing custom GPU kernels for attention, mixture-of-experts, quantization, and operator fusion, integrating improvements into vLLM, SGLang, and PyTorch, and collaborating with architecture and compiler teams on GPU roadmaps. Applicants need a bachelor's degree plus 4 or more years of experience, a master's plus 3 or more years, or a PhD, along with relevant GPU computing, AI systems, or high-performance computing experience and proficiency in modern C++ and Python. The position is hybrid and may be based in Santa Clara or other listed Intel locations in California, Oregon, or Texas.
