Description
The AI Platform role operates and improves large-scale LLM and embedding model-serving frameworks such as vLLM, Dynamo, and Triton; builds and monitors a reliable real-time serving environment; optimizes resource efficiency, inference speed, and scalability; and supports incident response, deployment stability, orchestration, and automation. The role seeks experience with GPU and LLM internals, serving frameworks, operational stability, and high-performance serving, with additional value from generalized AI/ML platform experience.
