Description
Plaud is hiring a SpeechLLM Research Engineer to build and optimize high-throughput, low-latency inference systems for large language models and foundational speech models. The role focuses on real-time streaming, continuous batching, KV cache management, GPU architecture, distributed multi-GPU inference, and model compression or quantization. Candidates should have hands-on experience with LLM or speech-model serving, GPU systems, and real-time conversational AI. The position offers a $195,000–$365,000 base salary plus bonus and equity, healthcare, retirement matching, paid time off, and a hybrid office arrangement requiring at least three days in the office per week.

