Description
NVIDIA is hiring a Solutions Architect focused on AI inference to build and optimize distributed inference pipelines using NVIDIA Dynamo, Kubernetes, TensorRT-LLM, vLLM, SGLang, and related technologies. The role involves collaborating with engineering, DevOps, and customers; deploying disaggregated inference systems; providing technical leadership; and resolving complex GPU, memory, and networking issues. Candidates need at least five years of solutions architecture experience, relevant NVIDIA inference and GPU orchestration expertise, and a BS in computer science or engineering or equivalent experience. The base salary range is 152,000–241,500 USD for Level 3 and 184,000–287,500 USD for Level 4, with equity and benefits available.
