Description
CEA is hiring a six-month internship in Saclay, France, focused on developing a quantization-aware and hardware-aware token pruning method for Vision Transformers deployed on resource-constrained embedded systems. The intern will characterize quantization effects on token selection, co-optimize pruning and quantization, evaluate pruning granularity and tensor structure, and integrate the methods into Aidge, a deep-learning platform for embedded neural networks. The work involves Python, PyTorch, Linux, C++, and deep learning, with validation against accuracy, latency, memory footprint, and pruning overhead.
