Summary from listing
The Inference Engineer will build and optimize high-performance inference systems for Large Language Models, Speech AI, and multimodal AI workloads. The role focuses on low latency, high throughput, GPU efficiency, scalable serving infrastructure, distributed inference, and cost optimization, while collaborating with ML researchers, platform engineers, speech AI teams, and product engineering teams.
