Description
Qualcomm is hiring LLM Serving Engineers to build scalable LLM inference platforms and contribute to serving packages such as vLLM, SGLang, Triton-Inference Server, and others. The role covers inference techniques, model optimization, autoscaling, load balancing, routing, customer collaboration, and open-source contributions. Candidates need hands-on experience with LLM serving or orchestration packages, strong Python and distributed-systems skills, deep understanding of transformer-based models, and a relevant bachelor's, master's, or PhD degree with the specified years of software engineering experience.
