Description
Intel is hiring a college-graduate engineer for its Neural Compressor team in Shanghai, China. The role develops and optimizes model compression tools such as Intel Neural Compressor and AutoRound for Intel CPUs, GPUs, and AI accelerators; researches quantization and compression for large language, vision-language, and generative models; and explores efficient deployment, inference acceleration, and fine-tuning acceleration. The position is on-site and requires a computer science-related bachelor's or master's degree, deep learning and LLM knowledge, familiarity with quantization and pruning, and programming proficiency in Python or C++.
