Description
DRW is hiring an AI Inference Platform Engineer to build, operate, and optimize production systems serving large language, vision, multimodal, and embedding models. The role owns end-to-end inference platform engineering, including NVIDIA GPU serving, model onboarding, performance profiling, distributed and multi-tenant scheduling, reliability, observability, and cost optimization. Candidates need hands-on LLM serving experience, expertise in modern inference runtimes, knowledge of GPU architectures and optimization techniques, and strong Linux and systems performance skills. The annual base salary is $200,000 to $250,000, with a discretionary bonus and comprehensive benefits.
