Description
Qualcomm is hiring a Staff Engineer – AI Model Optimization Architect in Cork to lead end-to-end model transformation and optimization for large language models, vision-language models, diffusion models, and multimodal models on Qualcomm inference accelerators. The role covers PyTorch and ONNX model rewrites, graph capture, Triton-based fusion kernels, transformer and KV-cache optimizations, continuous batching, distributed inference, and scalable production deployment. Candidates need expert PyTorch and inference optimization expertise, strong Python and compiler experience, transformer and accelerator knowledge, and a relevant bachelor’s, master’s, or PhD degree with the stated years of software engineering experience.
