Description
Qualcomm is hiring LLM Serving Engineers to build scalable LLM inference platforms and contribute to serving packages such as vLLM, SGLang, Triton-Inference Server, and others. The role covers inference techniques, model optimization, autoscaling, load balancing, routing, customer collaboration, and open-source contributions. Candidates need strong Python, PyTorch, distributed systems, computer architecture, and deep-learning optimization experience, along with a relevant bachelor's, master's, or PhD degree and the specified years of related work experience. The listed pay range is $162,000 to $243,000.
